Jobs / Roche

Principal/Senior Site Reliability Engineer

Roche · London, ENG, United Kingdom
London, ENG, United KingdomRemote
Remuneration
Not specified
Location
London, ENG, United Kingdom
Visa sponsorship
Not specified

Job summary

Join Roche as a Senior Site Reliability Engineer in the Computational Sciences Center of Excellence, where you will design resilient, cloud-based systems for MLOps and HPC workloads.

Benefits

Relocation benefits

Qualifications

  • expertise in Infrastructure as Code
  • understand cloud-native and on-prem architectures
  • hands-on with Docker, Kubernetes, and Kubeflow
  • expert in automation scripting
  • degree in Computer Science or related field

Responsibilities

  • architect Infrastructure as Code using Terraform, Pulumi, or CloudFormation
  • design disaster recovery and failover plans
  • strengthen reliability through chaos engineering
  • build observability with monitoring, logging, and alerting frameworks
  • provide technical leadership to a team of engineers
  • align infrastructure with ML and HPC needs

Skills

AirflowAWSAzureBashCloudFormationDatadogDockerGCPGoGrafanaKubernetesPrometheusPulumiPythonSparkTerraform

Degrees

Computer Science

Relocation

No