Jobs / IQVIA
Senior AI Platform Engineer
IQVIA · London, ENG, United Kingdom
London, ENG, United KingdomHybrid
Remuneration
Not specified
Location
London, ENG, United Kingdom
Visa sponsorship
Not specified
Job summary
The Senior AI Platform Engineer at IQVIA is responsible for defining and delivering the infrastructure strategy for Large Language Model (LLM) programmes, providing technical leadership across various teams. This role involves designing and operating platforms for training, evaluating, deploying, and governing large-scale AI systems in healthcare use cases.
Qualifications
- Significant experience designing, building, and operating large-scale AI, machine learning, or distributed computing platforms in enterprise environments.
- Deep understanding of LLM architectures and their interaction with GPU infrastructure, including CUDA, cuDNN, NCCL, kernel-level acceleration libraries, and distributed training frameworks such as PyTorch.
- Strong knowledge of distributed training and inference strategies, including tensor, pipeline, data, and expert parallelism approaches.
- Experience optimising LLM inference workloads using technologies such as vLLM, TensorRT-LLM, NVIDIA NIM, SGLang, or similar high-performance serving frameworks.
- Expertise in model optimisation techniques including quantisation, mixed precision training and inference (FP8, GPTQ, AWQ, LoRA), and performance tuning for large-scale model deployment.
- Advanced experience profiling, troubleshooting, and optimising GPU workloads using tools such as NVIDIA Nsight, DCGM, and related ecosystem technologies.
- Strong background in AWS cloud services, high-performance computing, distributed systems, containerised environments, and infrastructure automation.
- Experience with workload orchestration technologies such as Slurm, Kubernetes, Ray, or equivalent distributed compute frameworks.
- Demonstrated success bridging research and production environments, enabling rapid experimentation while maintaining operational excellence, governance, security, and reliability.
- Proven ability to lead complex cross-functional initiatives, influence technical direction, and communicate effectively with engineering, research, product, and executive stakeholders.
Responsibilities
- Own the AI platform and infrastructure roadmap, leading the planning and execution of LLM initiatives and translating research requirements into scalable engineering solutions.
- Partner with centralised infrastructure teams to design and deliver high-performance compute environments across AWS and on-premises platforms, including GPU infrastructure, Slurm clusters, and migration from ad hoc research workflows.
- Optimise LLM training and inference workloads, supporting research and product teams in maximising performance, scalability, and reliability across the infrastructure stack.
- Establish and maintain model and data lifecycle capabilities, including dataset versioning, lineage tracking, reproducibility standards, and integration with model registries.
- Lead the evolution of knowledge graph infrastructure, driving technology selection, migration strategies, performance optimisation, and integration with AI workflows.
- Serve as the primary technical coordination point across AI Research, Data Engineering, MLOps, Product, and Infrastructure teams, resolving dependencies and prioritising activities critical to delivery.
- Provide technical leadership for vendor selection, procurement, and technology partnerships, advising on compute architectures, GPU specifications, AI platforms, and integration approaches.
- Define platform engineering standards, governance, and best practices while mentoring engineers and promoting operational excellence across AI and platform teams.
Skills
AWSKubernetes
Relocation
No