Jobs / CloudFactory

Senior Site Reliability Engineer

CloudFactory · Berlin, BE, Deutschland
Berlin, BE, DeutschlandExp: 5+ yrsOnsite
Remuneration
Not specified
Location
Berlin, BE, Deutschland
Visa sponsorship
Not specified

Job summary

The Site Reliability Engineer role is essential for ensuring the reliability and performance of production systems, particularly for machine learning and large language model workloads. Responsibilities include managing model serving infrastructure, enhancing observability for ML models, and developing reusable developer tooling and automation using technologies such as Kubernetes, Helm, and Terraform.

Qualifications

  • 5+ years in infrastructure engineering, DevOps, or SRE with large-scale production systems using Kubernetes
  • Production Operational experience
  • Proficiency in Python or Go for automation and tooling
  • Experience with AI tools in daily operations
  • First-principles reasoning
  • Ownership of at least one end-to-end infrastructure build
  • Cross-functional collaboration experience

Responsibilities

  • Reliability of platform including ML and LLM workloads
  • Observability for ML models
  • Company-wide technical direction
  • Developer tooling and automation
  • Reusable components packaging open-source tools
  • Secure-by-default infrastructure

Skills

AWSCloudFormationGitHubGitHub ActionsGoGrafanaHelmIstioKubernetesPythonTerraform

Relocation

No