Jobs / Tes***
Staff Site Reliability Engineer, Infrastructure Engineering
Tes*** · Austin, TX, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Austin, TX, United StatesRemote
Remuneration
Not specified
Location
Austin, TX, United States
Visa sponsorship
Sponsors visa
Job summary
Manage and develop the internal platform for AI/ML at Tes***, focusing on Kubernetes, GPU infrastructure, and modern ML platforms.
Qualifications
- Built agent sandbox/ephemeral compute platforms
- Built GPU sandboxes with secure isolated environments
- Configured and troubleshot RoCE v2 / RDMA fabrics
- Production experience with KServe, Triton, vLLM, or Ray Serve
- Built or operated training-as-a-service and job scheduling
- Integrated platform services into an internal cloud
- Root cause analysis and systems thinking
- Deep Kubernetes internals expertise
- Built production Kubernetes operators in Go
- Production Cluster API experience
Responsibilities
- Build and own the end-to-end AI/ML platform as a self-service product for internal users
- Write production Kubernetes operators and controllers in Go for GPU workloads and cluster lifecycle
- Operate large-scale GPU fleets scheduling and health monitoring
- Own RoCE/RDMA networking for distributed training
- Build and operate inference infrastructure using KServe, Triton, and Ray Serve
- Build and operate training-as-a-service and distributed training
- Design secure, isolated sandbox environments for AI agents
- Architect serverless ephemeral workloads and event-driven compute
- Integrate GPU compute and sandboxes as services in the internal cloud platform
- Manage Kubernetes cluster lifecycle at fleet scale using Cluster API
Skills
etcdGoIAMKnativeKubernetesAWS Lambda
Relocation
No