Jobs / Roche
Principal/Senior Site Reliability Engineer
Roche · London, ENG, United Kingdom
London, ENG, United KingdomRemote
Remuneration
Not specified
Location
London, ENG, United Kingdom
Visa sponsorship
Not specified
Job summary
Join Roche as a Senior Site Reliability Engineer in the Computational Sciences Center of Excellence, where you will design resilient, cloud-based systems for MLOps and HPC workloads.
Benefits
Relocation benefits
Qualifications
- expertise in Infrastructure as Code
- understand cloud-native and on-prem architectures
- hands-on with Docker, Kubernetes, and Kubeflow
- expert in automation scripting
- degree in Computer Science or related field
Responsibilities
- architect Infrastructure as Code using Terraform, Pulumi, or CloudFormation
- design disaster recovery and failover plans
- strengthen reliability through chaos engineering
- build observability with monitoring, logging, and alerting frameworks
- provide technical leadership to a team of engineers
- align infrastructure with ML and HPC needs
Skills
AirflowAWSAzureBashCloudFormationDatadogDockerGCPGoGrafanaKubernetesPrometheusPulumiPythonSparkTerraform
Degrees
Computer Science
Relocation
No