Jobs / Develocity
Staff Site Reliability Engineer
Develocity · United Kingdom · Remote
United KingdomExp: 7+ yrsRemote
Remuneration
Not specified
Location
United Kingdom · Remote
British Summer Time (UTC+1)
Visa sponsorship
Not specified
Job summary
Develocity is seeking a Lead Site Reliability Engineer to establish a new SRE team, focusing on reliability and performance of their toolchain observability platform.
Qualifications
- 7+ years in SRE, DevOps, or equivalent role operating production services.
- Experience leading reliability initiatives across teams.
- Ability to influence technical direction without direct authority.
- Experience with SLOs and error budgets.
- Strong Kubernetes experience in production.
- Cloud infrastructure expertise, preferably AWS.
- Proficiency with observability tools and Infrastructure as Code.
- Track record of incident management in a 24/7 environment.
- Scripting proficiency for automation.
- Strong written and verbal English communication skills.
Responsibilities
- Operate and maintain Develocity instances and supporting services in production.
- Define and evolve SRE standards, practices, and operating models.
- Participate in on-call rotation as a technical escalation point.
- Lead incident response and retrospectives for reliability improvements.
- Set reliability priorities based on risk and customer impact.
- Identify reliability risks and evolve SaaS operations.
- Lead architectural and design reviews for reliability and scalability.
- Drive automation across operational workflows.
- Build and maintain observability for managed services.
- Own disaster recovery and business continuity planning.
- Partner with engineering leadership for feature delivery and reliability.
- Mentor and coach SREs for technical growth.
- Onboard new SREs and contribute to hiring.
- Communicate with customers during incidents.
- Optimize performance and operational costs.
Skills
AWSBashEKSGradleGrafanaJavaKotlinKubernetesMavennpmPrometheusPythonS3TerraformWindows
Relocation
No