Jobs / Bit***

Senior AI Storage Infrastructure Engineer

Bit*** · San Jose, CA, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
San Jose, CA, United StatesExp: 5+ yrsRemote
Remuneration
Not specified
Location
San Jose, CA, United States
Visa sponsorship
Sponsors visa

Job summary

Seeking a Senior AI Storage Infrastructure Engineer to architect high-performance storage solutions for AI-native NeoCloud, focusing on I/O intensive AI model training and inference workloads.

Qualifications

  • 5+ years of experience in distributed storage systems and high-performance file systems
  • Deep expertise in Kubernetes CSI paradigm and building volume plugins
  • Strong experience with block/file I/O at the Linux OS level and performance tuning
  • Familiarity with high-throughput networking protocols and their interaction with storage subsystems
  • Proven track record of operating and scaling large-scale storage environments
  • Experience with infrastructure automation tools and CI/CD pipelines
  • Excellent technical communication skills
  • Experience in high-velocity, high-growth engineering environments preferred

Responsibilities

  • Design, deploy, and maintain Container Storage Interface (CSI) drivers for high-performance parallel file systems
  • Architect and implement GPUDirect Storage (GDS) integrations for direct memory access between NVMe drives and GPU memory
  • Develop and manage local NVMe caching strategies for low-latency loading of model weights and datasets
  • Optimize IOPS, throughput, and latency across the containerized storage stack
  • Collaborate with GPU Systems & Fabric team to optimize storage layer for RDMA and high-speed interconnects
  • Implement automated monitoring and alerting for storage performance
  • Define storage policies, quota management, and multi-tenancy strategies within Kubernetes
  • Mentor junior engineers and drive architectural design reviews

Skills

AnsibleKubernetesLinuxTerraform

Relocation

No