AI Infrastructure Engineer
You will architect, deploy, and maintain infrastructure for large-scale AI compute environments in a high-density AI data center. You will manage GPU clusters, optimize networking, maintain high-throughput storage, automate provisioning, and troubleshoot compute, networking, and storage performance.
Responsibilities
- Deploy and manage large-scale GPU clusters using Kubernetes or Slurm
- Optimize high-speed low-latency networking for distributed compute
- Plan and monitor rack density across AI infrastructure
- Implement and maintain high-throughput storage systems for GPU-intensive workloads
- Automate infrastructure provisioning and configuration using Infrastructure as Code tools
- Troubleshoot and optimize compute networking and storage performance
Requirements
- Degree in Computer Science Data Engineering or a related technical field
- Strong experience with Linux administration containerization and GPU infrastructure
- Experience with Kubernetes or Slurm
- Familiarity with CUDA NCCL and Triton Inference Server
- Understanding of NVIDIA GB300 and VR NVL72 Scalable Units
- Experience in HPC AI infrastructure or large-scale distributed compute environments
- Experience with InfiniBand RoCE v2 or high-performance networking architectures
- Familiarity with Lustre BeeGFS or WekaIO
- Experience with Terraform Ansible or other Infrastructure as Code frameworks
- NVIDIA Kubernetes or cloud infrastructure certifications