Staff Cloud Support Engineer
You serve as the highest-level escalation point for complex incidents, lead cross-functional root-cause investigations, and design systemic reliability improvements. You influence Kubernetes and workload orchestration architecture, troubleshoot AI and machine-learning infrastructure, advise customers during high-risk incidents, deliver executive-ready root-cause analyses, mentor engineers, and define support standards.
Responsibilities
- Serve as the highest-level escalation point for complex P1/P0 incidents
- Lead cross-functional root-cause investigations
- Design systemic fixes with SRE and software teams
- Improve node validation, burn-in, performance baselining, and release readiness
- Influence Kubernetes architecture and workload orchestration
- Reduce MTTR and incident recurrence
- Troubleshoot NCCL, Infiniband, GPU driver, and firmware issues
- Support AI training and inference workloads
- Deliver executive-ready root-cause analyses
- Mentor engineers and define technical standards
Requirements
- 8+ years of experience in SRE, DevOps, HPC, or Cloud Infrastructure roles
- Advanced Linux systems expertise
- Deep Kubernetes operational experience at CKA level or higher
- Strong knowledge of Infiniband, RDMA, RoCE, and SDN
- Experience supporting AI/ML workloads at scale on GPU clusters
- Track record of resolving multi-layer distributed system failures
- Strong customer communication and executive-facing presence
Benefits
- Restricted Stock Units
- Paid time off
- Paid holidays
- Comprehensive health insurance
- Dental insurance
- Vision insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance
- Short-term disability insurance
- Long-term disability insurance
- Professional development
- Tuition reimbursement
- Mental health and wellness support
- Commuter benefits
- Cell phone stipend
- 401(k) retirement plan with company match up to 4% of salary
- Volunteer time off