Jobs / CloudFactory
Senior Site Reliability Engineer
CloudFactory · Berlin, BE, Deutschland
Berlin, BE, DeutschlandExp: 5+ yrsOnsite
Remuneration
Not specified
Location
Berlin, BE, Deutschland
Visa sponsorship
Not specified
Job summary
The Site Reliability Engineer role is essential for ensuring the reliability and performance of production systems, particularly for machine learning and large language model workloads. Responsibilities include managing model serving infrastructure, enhancing observability for ML models, and developing reusable developer tooling and automation using technologies such as Kubernetes, Helm, and Terraform.
Qualifications
- 5+ years in infrastructure engineering, DevOps, or SRE with large-scale production systems using Kubernetes
- Production Operational experience
- Proficiency in Python or Go for automation and tooling
- Experience with AI tools in daily operations
- First-principles reasoning
- Ownership of at least one end-to-end infrastructure build
- Cross-functional collaboration experience
Responsibilities
- Reliability of platform including ML and LLM workloads
- Observability for ML models
- Company-wide technical direction
- Developer tooling and automation
- Reusable components packaging open-source tools
- Secure-by-default infrastructure
Skills
AWSCloudFormationGitHubGitHub ActionsGoGrafanaHelmIstioKubernetesPythonTerraform
Relocation
No