Jobs / Ana***
ML Ops Engineer
Ana*** · London, ENG, United Kingdom
Visa sponsorship details are locked. Unlock company name and apply link with .
London, ENG, United KingdomOnsite
Remuneration
Not specified
Location
London, ENG, United Kingdom
Visa sponsorship
Sponsors visa
Job summary
Ana*** is seeking a ML Ops Engineer to join their Platform Engineering team, focusing on designing, scaling, and maintaining high-performance MLOps and LLMOps infrastructure.
Qualifications
- Hands-on production experience in DevOps, SRE, or Platform Engineering
- Proven track record of deploying and operationalising machine learning models
- Experience managing compute-intensive GPU infrastructure
- Advanced proficiency in Kubernetes, Docker, Helm, KubeFlow
- Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, Jenkins
- Experience with vLLM, Ray, MLflow, LangChain, DeepSpeed, Hugging Face TGI
- Solid background in AWS, GCP, Azure, and GPU cost optimisation techniques
- Strong skills in Python, Bash, or Go; knowledge of Linux kernel tuning
Responsibilities
- Provision and manage cloud-native AI/ML infrastructure
- Automate core platform infrastructure using Infrastructure as Code tools
- Optimise GPU compute workloads and storage
- Build and maintain CI/CD and MLOps pipelines
- Deploy Large Language Models and generative AI workloads
- Enable automated model validation and monitoring
- Monitor and optimise cloud spend across GPU/CPU clusters
- Implement auto-scaling strategies and dynamic resource allocation
- Establish benchmarking and telemetry for AI models
- Implement end-to-end observability using various tools
Skills
AnsibleArgo CDAWSAzureBashDockerGCPGitHubGitHub ActionsGoGrafanaHelmIstioJenkinsKubernetesLinuxOpenTelemetryPrometheusPythonTerraform
Relocation
No