Jobs / Ana***

ML Ops Engineer

Ana*** · London, ENG, United Kingdom
Visa sponsorship details are locked. Unlock company name and apply link with .
London, ENG, United KingdomOnsite
Remuneration
Not specified
Location
London, ENG, United Kingdom
Visa sponsorship
Sponsors visa

Job summary

Ana*** is seeking a ML Ops Engineer to join their Platform Engineering team, focusing on designing, scaling, and maintaining high-performance MLOps and LLMOps infrastructure.

Qualifications

  • Hands-on production experience in DevOps, SRE, or Platform Engineering
  • Proven track record of deploying and operationalising machine learning models
  • Experience managing compute-intensive GPU infrastructure
  • Advanced proficiency in Kubernetes, Docker, Helm, KubeFlow
  • Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, Jenkins
  • Experience with vLLM, Ray, MLflow, LangChain, DeepSpeed, Hugging Face TGI
  • Solid background in AWS, GCP, Azure, and GPU cost optimisation techniques
  • Strong skills in Python, Bash, or Go; knowledge of Linux kernel tuning

Responsibilities

  • Provision and manage cloud-native AI/ML infrastructure
  • Automate core platform infrastructure using Infrastructure as Code tools
  • Optimise GPU compute workloads and storage
  • Build and maintain CI/CD and MLOps pipelines
  • Deploy Large Language Models and generative AI workloads
  • Enable automated model validation and monitoring
  • Monitor and optimise cloud spend across GPU/CPU clusters
  • Implement auto-scaling strategies and dynamic resource allocation
  • Establish benchmarking and telemetry for AI models
  • Implement end-to-end observability using various tools

Skills

AnsibleArgo CDAWSAzureBashDockerGCPGitHubGitHub ActionsGoGrafanaHelmIstioJenkinsKubernetesLinuxOpenTelemetryPrometheusPythonTerraform

Relocation

No