Jobs / JPMorganChase

Lead SRE - AWS,Python

JPMorganChase · Glasgow, SCT, United Kingdom
Glasgow, SCT, United KingdomHybrid
Remuneration
Not specified
Location
Glasgow, SCT, United Kingdom
Visa sponsorship
Not specified

Job summary

As a Lead Site Reliability Engineer, you will ensure the availability, performance, and resilience of production systems serving millions globally. Key responsibilities include leading the design of scalable infrastructure solutions, driving incident response and root cause analysis, and developing automation frameworks to enhance deployment pipelines, utilizing technologies such as Python, Kubernetes, and Terraform.

Qualifications

  • Hands-on experience designing and operating large-scale distributed systems
  • Proficiency in one or more programming or scripting languages
  • Strong background in observability tooling
  • Demonstrated experience leading incident response processes
  • Experience with container orchestration and infrastructure-as-code practices

Responsibilities

  • Lead the design and implementation of scalable, reliable, and observable infrastructure solutions
  • Define and enforce service level objectives, error budgets, and reliability targets
  • Drive incident response, root cause analysis, and post-incident reviews
  • Develop and maintain automation frameworks to eliminate toil
  • Collaborate cross-functionally with software engineering, architecture, and security teams
  • Champion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions
  • Mentor and guide junior engineers
  • Evaluate and influence platform and tooling decisions
  • Leverage enterprise-authorized AI-assisted engineering practices

Skills

AWSBashGoJavaKubernetesPythonTerraform

Certifications

Formal training or certification on site reliability engineering concepts

Relocation

No