AI Infrastructure Engineer

You will architect, deploy, and maintain infrastructure for large-scale AI compute environments in a high-density AI data center. You will manage GPU clusters, optimize networking, maintain high-throughput storage, automate provisioning, and troubleshoot compute, networking, and storage performance.

Responsibilities

  • Deploy and manage large-scale GPU clusters using Kubernetes or Slurm
  • Optimize high-speed low-latency networking for distributed compute
  • Plan and monitor rack density across AI infrastructure
  • Implement and maintain high-throughput storage systems for GPU-intensive workloads
  • Automate infrastructure provisioning and configuration using Infrastructure as Code tools
  • Troubleshoot and optimize compute networking and storage performance

Requirements

  • Degree in Computer Science Data Engineering or a related technical field
  • Strong experience with Linux administration containerization and GPU infrastructure
  • Experience with Kubernetes or Slurm
  • Familiarity with CUDA NCCL and Triton Inference Server
  • Understanding of NVIDIA GB300 and VR NVL72 Scalable Units
  • Experience in HPC AI infrastructure or large-scale distributed compute environments
  • Experience with InfiniBand RoCE v2 or high-performance networking architectures
  • Familiarity with Lustre BeeGFS or WekaIO
  • Experience with Terraform Ansible or other Infrastructure as Code frameworks
  • NVIDIA Kubernetes or cloud infrastructure certifications

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available