Remuneration
Not specified
Location
Deutschland
Visa sponsorship
Not specified
Job summary
Site Reliability Engineer at Runware ensuring reliability, performance, and resilience of production services.
Benefits
Generous paid time offMeaningful stock optionsRemote-first setupFlexible hoursFamily leaveCompany retreats
Qualifications
- Experience operating and troubleshooting production systems at scale
- Strong understanding of distributed systems and debugging across various components
- Experience designing and operating observability systems
- Understanding of SRE principles including SLIs, SLOs, and incident management
- Experience with Kubernetes, containers, IaC, and automated deployment practices
- Ownership of production problems and participation in on-call rotation
Responsibilities
- Improve reliability, availability, and performance of production services
- Define and evolve reliability practices including SLIs, SLOs, and observability standards
- Investigate production issues across distributed systems and participate in on-call rotation
- Lead incident reviews and turn failure modes into engineering improvements
- Reduce operational toil through automation and system resilience improvements
- Collaborate with Engineering and DevOps on capacity planning and architectural improvements
Skills
ClickHouseGoKubernetesMySQLPHPPythonRabbitMQRedis
Relocation
No