Jobs / CloudFactory
Senior Site Reliability Engineer
CloudFactory · Reading, ENG, United Kingdom
Reading, ENG, United KingdomExp: 5+ yrsOnsite
Remuneration
Not specified
Location
Reading, ENG, United Kingdom
Visa sponsorship
Not specified
Job summary
The Site Reliability Engineer role is essential for ensuring the reliability and performance of production systems, particularly for machine learning and large language model workloads. Responsibilities include managing model serving infrastructure, enhancing observability for ML models, and developing automation tools using technologies such as Kubernetes, Helm, and Terraform to streamline operations and improve security.
Qualifications
- 5+ years in infrastructure engineering, DevOps, or SRE
- Production Operational experience
- Good proficiency in Python or Go
- AI is already in your daily loop
- First-principles reasoning
- At least one infrastructure build you owned end to end
- Cross-functional strength
Responsibilities
- Reliability of platform (includes ML and LLM workloads)
- Observability (includes ML models)
- Company-wide technical direction
- Developer tooling and automation
- Reusable components packaging common open-source tools
- Secure-by-default infrastructure
Skills
AWSCloudFormationGitHubGitHub ActionsGoGrafanaHelmIstioKubernetesPythonTerraform
Relocation
No