Jobs / Scalingo

Lead Site Reliability Engineer - Cloud - Remote - F/H

Scalingo · Strasbourg, GE, France · Remote
Strasbourg, GE, France60,000-70,000 EUR/yearlyRemote
Remuneration
60,000-70,000 EUR/yearly
Location
Strasbourg, GE, France · Remote
Central European Summer Time (UTC+2)
Visa sponsorship
Not specified

Job summary

Lead Site Reliability Engineer responsible for ensuring platform reliability and performance.

Qualifications

  • Strong expertise in cloud environments and distributed infrastructures with a focus on high availability and reliability.
  • Mastery of observability practices (logs, metrics, alerting) and structured diagnostic capabilities for complex incidents.
  • Good understanding of containerized environments and their operational challenges.
  • Proven skills in production databases: reliability, backups, restoration, replication, and scaling.
  • Experience with Infrastructure as Code and environment automation.
  • Sensitivity to operational security issues.
  • Proficiency in using AI tools for daily efficiency.
  • Ability to navigate complex, changing, or uncertain contexts with rigor and reliability.
  • Skill in prioritization, including during incidents.
  • Clear and structured communication, with a taste for cross-collaboration and knowledge sharing.
  • Blameless attitude, technical curiosity, composure, and user impact awareness.
  • Ability to exercise technical leadership and advance collective practices.

Responsibilities

  • Ensure stability, availability, and resilience of production systems.
  • Anticipate failures and structure effective incident responses.
  • Industrialize and automate platform operations.
  • Maintain high service quality for clients and contractual commitments.
  • Lead team organization (processes, rituals, documentation).
  • Support prioritization and review technical choices and implementations.
  • Facilitate team skill development, promoting autonomy and initiative.
  • Transmit SRE best practices (reliability, observability, incident management, automation).
  • Implement technical SRE vision in structuring projects.
  • Analyze performance, identify bottlenecks, and propose resource optimization.
  • Define, implement, and improve observability tools (monitoring, metrics, logs, alerting).
  • Document and evolve operational processes.
  • Conduct continuous technological monitoring for infrastructure improvements.
  • Provide level 3 client support in coordination with support teams.
  • Actively participate in incident management and on-call cycles.
  • Respond quickly to critical incidents to minimize impact.
  • Lead incident retrospectives, identifying root causes and defining corrective actions.
  • Draft and publish post-mortem reports after major incidents.
  • Coordinate and communicate during crises, both internally and with clients.
  • Ensure compliance with service commitments (SLA, RPO, RTO) in SRE scope.

Skills

ClairLinux

Relocation

No