Jobs / Qua***
Senior Platform Engineer
Qua*** · United States · Remote
Visa sponsorship details are locked. Unlock company name and apply link with .
United StatesExp: 10+ yrsRemote
Remuneration
Not specified
Location
United States · Remote
Eastern Daylight Time (UTC-4)
Visa sponsorship
Sponsors visa
Job summary
Qua*** is seeking a Senior Platform Engineer to design, optimize, and scale infrastructure for GenAI and LLM workloads. The role involves collaborating with cross-functional teams to deploy AI solutions and requires deep hands-on experience in GPU profiling and distributed training environments.
Qualifications
- Strong experience with Slurm and distributed training environments
- Hands-on expertise with Red Hat OpenShift and/or Kubernetes
- Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT)
- Strong foundation in Linux systems, performance tuning, and multi-GPU optimization
- Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems)
- Familiarity with Infrastructure-as-Code tools (Terraform, Ansible)
- Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters
Responsibilities
- Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments
- Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads
- Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments
- Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS)
- Collaborate with cross-functional teams to deploy models in research and production environments
- Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps)
- Develop reusable infrastructure templates using Terraform and Helm
- Contribute to internal innovation and support client-facing delivery engagements
Skills
AnsibleAWSAzureGCPHelmKubernetesLinuxMakeOpenShiftOracle CloudRHELSnowflakeTerraform
Relocation
No