Jobs / JPMorganChase
Lead SRE - AWS,Python
JPMorganChase · Glasgow, SCT, United Kingdom
Glasgow, SCT, United KingdomHybrid
Remuneration
Not specified
Location
Glasgow, SCT, United Kingdom
Visa sponsorship
Not specified
Job summary
As a Lead Site Reliability Engineer, you will ensure the availability, performance, and resilience of production systems serving millions globally. Key responsibilities include leading the design of scalable infrastructure solutions, driving incident response and root cause analysis, and developing automation frameworks to enhance deployment pipelines, utilizing technologies such as Python, Kubernetes, and Terraform.
Qualifications
- Hands-on experience designing and operating large-scale distributed systems
- Proficiency in one or more programming or scripting languages
- Strong background in observability tooling
- Demonstrated experience leading incident response processes
- Experience with container orchestration and infrastructure-as-code practices
Responsibilities
- Lead the design and implementation of scalable, reliable, and observable infrastructure solutions
- Define and enforce service level objectives, error budgets, and reliability targets
- Drive incident response, root cause analysis, and post-incident reviews
- Develop and maintain automation frameworks to eliminate toil
- Collaborate cross-functionally with software engineering, architecture, and security teams
- Champion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions
- Mentor and guide junior engineers
- Evaluate and influence platform and tooling decisions
- Leverage enterprise-authorized AI-assisted engineering practices
Skills
AWSBashGoJavaKubernetesPythonTerraform
Certifications
Formal training or certification on site reliability engineering concepts
Relocation
No