Member of Technical Staff - GPU Infrastructure

You will design, deploy, optimize, and support large-scale GPU infrastructure for customers. You will architect GPU clusters, implement orchestration and high-performance networking, configure parallel filesystems, tune system performance, resolve infrastructure issues across the stack, and provide operational documentation and support.

Responsibilities

  • Partner with clients to understand workload requirements
  • Design GPU cluster architectures
  • Create technical proposals and capacity plans
  • Develop deployment strategies for LLM training, inference, and HPC workloads
  • Present architectural recommendations
  • Deploy SLURM and Kubernetes
  • Implement InfiniBand, RoCE, and NVLink networking
  • Optimize GPU utilization and memory management
  • Configure Lustre, BeeGFS, and GPFS filesystems
  • Tune kernel and CUDA configurations
  • Resolve customer infrastructure issues
  • Implement monitoring, alerting, and automated remediation
  • Provide on-call support
  • Create runbooks and documentation

Requirements

  • 3+ years of hands-on experience with GPU clusters and HPC environments
  • SLURM and Kubernetes in production GPU settings
  • InfiniBand configuration and troubleshooting
  • NVIDIA GPU architecture, CUDA ecosystem, and driver stack
  • Ansible and Terraform
  • Python, Bash, and systems programming
  • Customer-facing technical leadership
  • NVIDIA drivers, Fabric Manager, and DCGM
  • Docker, Containerd, and Enroot
  • Linux kernel tuning
  • AI workload network topology
  • Power and cooling requirements for high-density GPU deployments

Benefits

  • Equity incentives

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available