Jobs / JPMorganChase
Senior Lead Software Engineer - LLM Ops Platform Reliability
JPMorganChase · Glasgow, SCT, United Kingdom
Glasgow, SCT, United KingdomOnsite
Remuneration
Not specified
Location
Glasgow, SCT, United Kingdom
Visa sponsorship
Not specified
Job summary
As a Senior Lead Software Engineer at JPMorganChase, you will build and operate large language model serving infrastructure, focusing on reliability, performance, and cost-efficiency. You will work with cloud and Kubernetes-based deployments, ensuring the stability and security of AI systems in production environments.
Qualifications
- Hands-on experience with system design, application development, testing, and operational stability in production environments
- Advanced proficiency in Python for building production-grade services and tooling
- Proficiency with automation and continuous delivery methods
- Hands-on experience with cloud infrastructure platforms and infrastructure-as-code tooling for delivery and lifecycle management
- Strong understanding of site reliability engineering practices, including incident management, root-cause analysis, runbooks, and reliability patterns
- Practical knowledge of observability and instrumentation across metrics, logs, and traces
- Hands-on experience with Kubernetes and container-based orchestration platforms, including managed cloud variants
- Experience hosting and serving large language models on cloud-based infrastructure and local GPU environments
- Knowledge of large language model reliability and risk considerations, including latency and throughput trade-offs, model versioning, prompt and response logging, and safe rollout patterns
- Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security
- Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations
Responsibilities
- Design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructure
- Build backend services and APIs that enable reliable operation of AI infrastructure in production environments
- Operate and scale large language model serving infrastructure, including model hosting, request routing, continuous batching, and cache optimization
- Deploy, host, and lifecycle-manage open-source and proprietary large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines
- Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads
- Tune GPU and accelerator capacity, autoscaling, and cost efficiency for large language model inference workloads using performance optimization techniques such as quantization, parallelism, and speculative decoding
- Lead reliability engineering for large language model endpoints through capacity planning, load and soak testing, safe rollouts, failover, and incident response for outages and model-quality regressions
- Participate in on-call rotations, lead incident triage and mitigation, and produce clear post-incident root-cause analyses and follow-up actions
- Identify recurring operational issues and automate remediation to improve platform stability and developer experience
- Build and maintain multi-agent systems with strong orchestration, including planning, coordination, tool-calling, state and memory management, and workflow control where appropriate
- Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes
Skills
KubernetesPythonAWSGCP
Relocation
No