Jobs / Bit***
Senior AI Storage Infrastructure Engineer
Bit*** · San Jose, CA, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
San Jose, CA, United StatesExp: 5+ yrsRemote
Remuneration
Not specified
Location
San Jose, CA, United States
Visa sponsorship
Sponsors visa
Job summary
Seeking a Senior AI Storage Infrastructure Engineer to architect high-performance storage solutions for AI-native NeoCloud, focusing on I/O intensive AI model training and inference workloads.
Qualifications
- 5+ years of experience in distributed storage systems and high-performance file systems
- Deep expertise in Kubernetes CSI paradigm and building volume plugins
- Strong experience with block/file I/O at the Linux OS level and performance tuning
- Familiarity with high-throughput networking protocols and their interaction with storage subsystems
- Proven track record of operating and scaling large-scale storage environments
- Experience with infrastructure automation tools and CI/CD pipelines
- Excellent technical communication skills
- Experience in high-velocity, high-growth engineering environments preferred
Responsibilities
- Design, deploy, and maintain Container Storage Interface (CSI) drivers for high-performance parallel file systems
- Architect and implement GPUDirect Storage (GDS) integrations for direct memory access between NVMe drives and GPU memory
- Develop and manage local NVMe caching strategies for low-latency loading of model weights and datasets
- Optimize IOPS, throughput, and latency across the containerized storage stack
- Collaborate with GPU Systems & Fabric team to optimize storage layer for RDMA and high-speed interconnects
- Implement automated monitoring and alerting for storage performance
- Define storage policies, quota management, and multi-tenancy strategies within Kubernetes
- Mentor junior engineers and drive architectural design reviews
Skills
AnsibleKubernetesLinuxTerraform
Relocation
No