Jobs / Lig***
Infrastructure Engineer (GPU & Compute)
Lig*** · London, ENG, United Kingdom
Visa sponsorship details are locked. Unlock company name and apply link with .
London, ENG, United KingdomExp: 5+ yrs180,000-200,000 GBP/yearlyHybrid
Remuneration
180,000-200,000 GBP/yearly
Location
London, ENG, United Kingdom
Visa sponsorship
Sponsors visa
Job summary
Lig*** is seeking a GPU & Compute Infrastructure Engineer to manage image management, system diagnostics, and validation across large-scale bare-metal compute infrastructure, focusing on GPU-enabled systems.
Benefits
Comprehensive medical, dental and vision coverage (U.S.); Private medical and deRetirement and financial wellness support (U.S.); Pension contribution (U.K.)Generous paid time off, plus holidaysPaid parental leaveProfessional development supportWellness and work-from-home stipendsFlexible work environment
Qualifications
- Experience with high-performance interconnects (e.g., InfiniBand, NVLink)
- Experience with PXE boot environments or image-based provisioning workflows
- Experience with hardware management interfaces such as iDRAC, IPMI, or Redfish
- Data center operations experience with physical hardware
- Experience supporting AI/ML or HPC workloads at scale
- Experience with GPU validation frameworks or large-scale hardware qualification processes
Responsibilities
- Own and evolve systems for image management, deployment, and validation across bare-metal infrastructure
- Run and maintain test clusters for system validation and diagnostics
- Validate firmware, drivers, and OS images across compute and GPU-enabled systems
- Support hardware qualification for next-generation platforms
- Own GPU diagnostics and validation workflows across large-scale infrastructure
- Diagnose and resolve complex issues across GPUs, drivers, OS, and hardware layers
- Analyze system and GPU performance using tools such as NVIDIA DCGM
- Identify failure patterns and improve system stability and validation coverage
- Build and maintain automation for provisioning and system bring-up
- Develop Python-based tools to improve efficiency and reduce manual overhead
- Improve reliability and scalability of image pipelines and validation systems
- Manage and operate Linux-based systems in production and validation environments
- Manage virtualization technology
- Support bare-metal provisioning workflows, including PXE and image-based systems
- Interface with hardware management systems for monitoring and debugging
- Collaborate with Infrastructure, Hardware, and Data Center teams on system bring-up and validation
- Work with platform and ML teams to ensure systems meet workload requirements
- Contribute to best practices for provisioning and lifecycle management of infrastructure
Skills
LinuxPython
Relocation
No