Jobs / NVIDIA
Senior Site Reliability Engineer, DGX Cloud
NVIDIA · Zürich, ZH, Switzerland
Zürich, ZH, SwitzerlandExp: 10+ yrsOnsite
Remuneration
Not specified
Location
Zürich, ZH, Switzerland
Visa sponsorship
Not specified
Job summary
NVIDIA is seeking a Senior Site Reliability Engineer for its DGX Cloud team to maintain high-performance clusters for AI researchers and enterprise clients.
Qualifications
- Expert knowledge of Kubernetes administration and microservices
- Experience with infrastructure automation tools
- Proficiency in a high-level programming language
- In-depth knowledge of Linux, networking, and cloud security
- Proficient in SRE principles and incident handling
- Experience with observability stacks and monitoring tools
Responsibilities
- Build and support large-scale Kubernetes clusters
- Define SLOs/SLIs and monitor error budgets
- Support services before launch with consulting and tools
- Maintain live services by monitoring availability and latency
- Optimize GPU workloads across various cloud platforms
- Scale systems through automation and improve reliability
- Lead triage and root-cause analysis of incidents
- Practice incident response and blameless postmortems
- Participate in on-call rotation for production services
Skills
AirflowAnsibleArgo WorkflowsAWSAzureChefGCPGoGrafanaKubernetesLightstepLinuxOpenTelemetryOracle CloudPrometheusPuppetPythonSplunkTerraform
Degrees
BS in Computer Science or related technical field
Relocation
No