Senior Staff Software Engineer DC Infrastructure
You will develop software that manages GPU servers and data centers, focusing on diagnostics, observability, automation, repair, and reliability for high-performance GPU clusters. You will build tooling and AI agents for hardware diagnosis, remediation, validation, facilities management, power, and liquid cooling, while owning deployment and operational support.
Responsibilities
- Develop deep-level diagnostics for GPU hardware faults
- Build troubleshooting and automation tooling for GPU platforms
- Develop AI agents for component diagnosis and hardware remediation
- Develop tooling for critical-environment management
- Build post-repair validation and testing tools
- Own deployment, monitoring, and operational support of tooling
- Develop facilities-management automation for power and liquid-cooling systems
Requirements
- Software engineering experience
- Ability to rapidly develop and ship scalable solutions
- Expertise in distributed systems, reliability, and cloud platforms
- Strength in Go, Python, Java, or Rust
- Experience with Kubernetes, infrastructure as code, and GCP
- Strong analytical and problem-solving skills
- Experience with Temporal and Kubernetes preferred
- Experience with large-scale GPU fleet operations or hyperscale data centers preferred
Benefits
- Industry competitive pay
- Restricted Stock Units
- Health insurance with HDHP and PPO options
- Vision insurance
- Dental insurance
- Employer contributions to HSA accounts
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability
- Teladoc
- 401(k) with 100% match up to 4% of salary
- Paid time off
- Paid holidays
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit of $300 per month