Jobs / CloudFactory

Senior Site Reliability Engineer

CloudFactory · Reading, ENG, United Kingdom
Reading, ENG, United KingdomExp: 5+ yrsOnsite
Remuneration
Not specified
Location
Reading, ENG, United Kingdom
Visa sponsorship
Not specified

Job summary

The Site Reliability Engineer role is essential for ensuring the reliability and performance of production systems, particularly for machine learning and large language model workloads. Responsibilities include managing model serving infrastructure, enhancing observability for ML models, and developing automation tools using technologies such as Kubernetes, Helm, and Terraform to streamline operations and improve security.

Qualifications

  • 5+ years in infrastructure engineering, DevOps, or SRE
  • Production Operational experience
  • Good proficiency in Python or Go
  • AI is already in your daily loop
  • First-principles reasoning
  • At least one infrastructure build you owned end to end
  • Cross-functional strength

Responsibilities

  • Reliability of platform (includes ML and LLM workloads)
  • Observability (includes ML models)
  • Company-wide technical direction
  • Developer tooling and automation
  • Reusable components packaging common open-source tools
  • Secure-by-default infrastructure

Skills

AWSCloudFormationGitHubGitHub ActionsGoGrafanaHelmIstioKubernetesPythonTerraform

Relocation

No