Jobs / NVIDIA

Senior Site Reliability Engineer, DGX Cloud

NVIDIA · Zürich, ZH, Switzerland
Zürich, ZH, SwitzerlandExp: 10+ yrsOnsite
Remuneration
Not specified
Location
Zürich, ZH, Switzerland
Visa sponsorship
Not specified

Job summary

NVIDIA is seeking a Senior Site Reliability Engineer for its DGX Cloud team to maintain high-performance clusters for AI researchers and enterprise clients.

Qualifications

  • Expert knowledge of Kubernetes administration and microservices
  • Experience with infrastructure automation tools
  • Proficiency in a high-level programming language
  • In-depth knowledge of Linux, networking, and cloud security
  • Proficient in SRE principles and incident handling
  • Experience with observability stacks and monitoring tools

Responsibilities

  • Build and support large-scale Kubernetes clusters
  • Define SLOs/SLIs and monitor error budgets
  • Support services before launch with consulting and tools
  • Maintain live services by monitoring availability and latency
  • Optimize GPU workloads across various cloud platforms
  • Scale systems through automation and improve reliability
  • Lead triage and root-cause analysis of incidents
  • Practice incident response and blameless postmortems
  • Participate in on-call rotation for production services

Skills

AirflowAnsibleArgo WorkflowsAWSAzureChefGCPGoGrafanaKubernetesLightstepLinuxOpenTelemetryOracle CloudPrometheusPuppetPythonSplunkTerraform

Degrees

BS in Computer Science or related technical field

Relocation

No