Senior Staff Software Engineer DC Infrastructure

You will develop software that manages GPU servers and data centers, focusing on diagnostics, observability, automation, repair, and reliability for high-performance GPU clusters. You will build tooling and AI agents for hardware diagnosis, remediation, validation, facilities management, power, and liquid cooling, while owning deployment and operational support.

Responsibilities

  • Develop deep-level diagnostics for GPU hardware faults
  • Build troubleshooting and automation tooling for GPU platforms
  • Develop AI agents for component diagnosis and hardware remediation
  • Develop tooling for critical-environment management
  • Build post-repair validation and testing tools
  • Own deployment, monitoring, and operational support of tooling
  • Develop facilities-management automation for power and liquid-cooling systems

Requirements

  • Software engineering experience
  • Ability to rapidly develop and ship scalable solutions
  • Expertise in distributed systems, reliability, and cloud platforms
  • Strength in Go, Python, Java, or Rust
  • Experience with Kubernetes, infrastructure as code, and GCP
  • Strong analytical and problem-solving skills
  • Experience with Temporal and Kubernetes preferred
  • Experience with large-scale GPU fleet operations or hyperscale data centers preferred

Benefits

  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance with HDHP and PPO options
  • Vision insurance
  • Dental insurance
  • Employer contributions to HSA accounts
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability
  • Teladoc
  • 401(k) with 100% match up to 4% of salary
  • Paid time off
  • Paid holidays
  • Cell phone reimbursement
  • Tuition reimbursement
  • Calm app subscription
  • MetLife Legal
  • Company-paid commuter benefit of $300 per month

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available