K8 Site Reliability SME
You design, deploy, and operate production Kubernetes control planes for GPU workloads. You manage GPU operators, topology-aware scheduling, workload lifecycle resources, tenant isolation, bare-metal provisioning, infrastructure as code, SLOs, monitoring, incident automation, and automated node failure recovery.
Responsibilities
- Operate production Kubernetes clusters for GPU workloads
- Configure the Nvidia GPU operator, device plugin, MIG, and GPU time slicing
- Implement topology-aware scheduling for GPU locality, NVLink domains, and network rails
- Develop CRDs for GPU workload lifecycle management
- Integrate Slurm, Ray, and Kubeflow with Kubernetes
- Implement multi-tenant isolation with namespaces, network policies, quotas, RBAC, and pod security
- Automate bare-metal provisioning, onboarding, lifecycle, and reclamation
- Build Terraform providers and modules for GPU cluster infrastructure
- Define and publish cluster availability, job completion, and provisioning latency SLOs
- Automate incident management and runbook execution
- Operate Prometheus, Grafana, Alertmanager, and PagerDuty monitoring
- Detect GPU node failures and automate drain, cordon, taint, and workload rescheduling
Requirements
- 5+ years in Kubernetes operations
- At least 2 years managing GPU workloads on Kubernetes
- Deep understanding of the Nvidia GPU operator, device plugin, and GPU scheduling
- Experience with topology-aware scheduling and GPU resource management
- Experience building multi-tenant Kubernetes platforms with strong isolation
- Experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom systems
- Proficiency in Terraform, Helm, and GitOps workflows such as ArgoCD or Flux
- Strong SRE background in SLI/SLO frameworks, incident management, and capacity planning
- Experience with Prometheus, Grafana, and alerting at scale
- Strong Go or Python programming skills for operator and CRD development
- Experience designing or implementing automated Kubernetes remediation
- Runbook-as-code mindset