Jobs / Lig***

Infrastructure Engineer (GPU & Compute)

Lig*** · London, ENG, United Kingdom
Visa sponsorship details are locked. Unlock company name and apply link with .
London, ENG, United KingdomExp: 5+ yrs180,000-200,000 GBP/yearlyHybrid
Remuneration
180,000-200,000 GBP/yearly
Location
London, ENG, United Kingdom
Visa sponsorship
Sponsors visa

Job summary

Lig*** is seeking a GPU & Compute Infrastructure Engineer to manage image management, system diagnostics, and validation across large-scale bare-metal compute infrastructure, focusing on GPU-enabled systems.

Benefits

Comprehensive medical, dental and vision coverage (U.S.); Private medical and deRetirement and financial wellness support (U.S.); Pension contribution (U.K.)Generous paid time off, plus holidaysPaid parental leaveProfessional development supportWellness and work-from-home stipendsFlexible work environment

Qualifications

  • Experience with high-performance interconnects (e.g., InfiniBand, NVLink)
  • Experience with PXE boot environments or image-based provisioning workflows
  • Experience with hardware management interfaces such as iDRAC, IPMI, or Redfish
  • Data center operations experience with physical hardware
  • Experience supporting AI/ML or HPC workloads at scale
  • Experience with GPU validation frameworks or large-scale hardware qualification processes

Responsibilities

  • Own and evolve systems for image management, deployment, and validation across bare-metal infrastructure
  • Run and maintain test clusters for system validation and diagnostics
  • Validate firmware, drivers, and OS images across compute and GPU-enabled systems
  • Support hardware qualification for next-generation platforms
  • Own GPU diagnostics and validation workflows across large-scale infrastructure
  • Diagnose and resolve complex issues across GPUs, drivers, OS, and hardware layers
  • Analyze system and GPU performance using tools such as NVIDIA DCGM
  • Identify failure patterns and improve system stability and validation coverage
  • Build and maintain automation for provisioning and system bring-up
  • Develop Python-based tools to improve efficiency and reduce manual overhead
  • Improve reliability and scalability of image pipelines and validation systems
  • Manage and operate Linux-based systems in production and validation environments
  • Manage virtualization technology
  • Support bare-metal provisioning workflows, including PXE and image-based systems
  • Interface with hardware management systems for monitoring and debugging
  • Collaborate with Infrastructure, Hardware, and Data Center teams on system bring-up and validation
  • Work with platform and ML teams to ensure systems meet workload requirements
  • Contribute to best practices for provisioning and lifecycle management of infrastructure

Skills

LinuxPython

Relocation

No