Jobs / Scalingo
Lead Site Reliability Engineer - Cloud - Remote - F/H
Scalingo · Strasbourg, GE, France · Remote
Strasbourg, GE, France60,000-70,000 EUR/yearlyRemote
Remuneration
60,000-70,000 EUR/yearly
Location
Strasbourg, GE, France · Remote
Central European Summer Time (UTC+2)
Visa sponsorship
Not specified
Job summary
Lead Site Reliability Engineer responsible for ensuring platform reliability and performance.
Qualifications
- Strong expertise in cloud environments and distributed infrastructures with a focus on high availability and reliability.
- Mastery of observability practices (logs, metrics, alerting) and structured diagnostic capabilities for complex incidents.
- Good understanding of containerized environments and their operational challenges.
- Proven skills in production databases: reliability, backups, restoration, replication, and scaling.
- Experience with Infrastructure as Code and environment automation.
- Sensitivity to operational security issues.
- Proficiency in using AI tools for daily efficiency.
- Ability to navigate complex, changing, or uncertain contexts with rigor and reliability.
- Skill in prioritization, including during incidents.
- Clear and structured communication, with a taste for cross-collaboration and knowledge sharing.
- Blameless attitude, technical curiosity, composure, and user impact awareness.
- Ability to exercise technical leadership and advance collective practices.
Responsibilities
- Ensure stability, availability, and resilience of production systems.
- Anticipate failures and structure effective incident responses.
- Industrialize and automate platform operations.
- Maintain high service quality for clients and contractual commitments.
- Lead team organization (processes, rituals, documentation).
- Support prioritization and review technical choices and implementations.
- Facilitate team skill development, promoting autonomy and initiative.
- Transmit SRE best practices (reliability, observability, incident management, automation).
- Implement technical SRE vision in structuring projects.
- Analyze performance, identify bottlenecks, and propose resource optimization.
- Define, implement, and improve observability tools (monitoring, metrics, logs, alerting).
- Document and evolve operational processes.
- Conduct continuous technological monitoring for infrastructure improvements.
- Provide level 3 client support in coordination with support teams.
- Actively participate in incident management and on-call cycles.
- Respond quickly to critical incidents to minimize impact.
- Lead incident retrospectives, identifying root causes and defining corrective actions.
- Draft and publish post-mortem reports after major incidents.
- Coordinate and communicate during crises, both internally and with clients.
- Ensure compliance with service commitments (SLA, RPO, RTO) in SRE scope.
Skills
ClairLinux
Relocation
No