Jobs / Tes***

Staff Site Reliability Engineer, Infrastructure Engineering

Tes*** · Austin, TX, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Austin, TX, United StatesRemote
Remuneration
Not specified
Location
Austin, TX, United States
Visa sponsorship
Sponsors visa

Job summary

Manage and develop the internal platform for AI/ML at Tes***, focusing on Kubernetes, GPU infrastructure, and modern ML platforms.

Qualifications

  • Built agent sandbox/ephemeral compute platforms
  • Built GPU sandboxes with secure isolated environments
  • Configured and troubleshot RoCE v2 / RDMA fabrics
  • Production experience with KServe, Triton, vLLM, or Ray Serve
  • Built or operated training-as-a-service and job scheduling
  • Integrated platform services into an internal cloud
  • Root cause analysis and systems thinking
  • Deep Kubernetes internals expertise
  • Built production Kubernetes operators in Go
  • Production Cluster API experience

Responsibilities

  • Build and own the end-to-end AI/ML platform as a self-service product for internal users
  • Write production Kubernetes operators and controllers in Go for GPU workloads and cluster lifecycle
  • Operate large-scale GPU fleets scheduling and health monitoring
  • Own RoCE/RDMA networking for distributed training
  • Build and operate inference infrastructure using KServe, Triton, and Ray Serve
  • Build and operate training-as-a-service and distributed training
  • Design secure, isolated sandbox environments for AI agents
  • Architect serverless ephemeral workloads and event-driven compute
  • Integrate GPU compute and sandboxes as services in the internal cloud platform
  • Manage Kubernetes cluster lifecycle at fleet scale using Cluster API

Skills

etcdGoIAMKnativeKubernetesAWS Lambda

Relocation

No