Senior AI Storage Infrastructure Engineer
You design, deploy, and maintain CSI drivers for high-performance parallel file systems; implement GPUDirect Storage integrations; manage local NVMe caching; optimize storage performance across the containerized stack; integrate storage with RDMA and high-speed interconnects; implement monitoring and alerting; define Kubernetes storage policies and multi-tenancy strategies; mentor junior engineers; and lead architectural reviews.
Responsibilities
- Design, deploy, and maintain CSI drivers for parallel file systems
- Architect and implement GPUDirect Storage integrations
- Develop and manage local NVMe caching strategies
- Optimize IOPS, throughput, and latency across the containerized storage stack
- Integrate storage with RDMA, InfiniBand, and RoCE
- Implement automated storage monitoring and alerting
- Define Kubernetes storage policies, quota management, and multi-tenancy isolation
- Mentor junior engineers
- Lead architectural design reviews
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field
- 5+ years of experience in distributed storage systems and high-performance file systems
- Deep understanding of POSIX compliance and file I/O semantics
- Expertise in the Kubernetes CSI paradigm
- Hands-on experience with block and file I/O at the Linux OS level
- Kernel-level performance tuning experience
- Familiarity with RDMA, InfiniBand, and RoCE
- Experience operating, debugging, and scaling large-scale storage environments
- Experience with Terraform, Ansible, and CI/CD pipelines
- Excellent technical communication skills
Benefits
- Attractive welfare benefits