K8 Site Reliability SME

You design, deploy, and operate production Kubernetes control planes for GPU workloads. You manage GPU operators, topology-aware scheduling, workload lifecycle resources, tenant isolation, bare-metal provisioning, infrastructure as code, SLOs, monitoring, incident automation, and automated node failure recovery.

Responsibilities

  • Operate production Kubernetes clusters for GPU workloads
  • Configure the Nvidia GPU operator, device plugin, MIG, and GPU time slicing
  • Implement topology-aware scheduling for GPU locality, NVLink domains, and network rails
  • Develop CRDs for GPU workload lifecycle management
  • Integrate Slurm, Ray, and Kubeflow with Kubernetes
  • Implement multi-tenant isolation with namespaces, network policies, quotas, RBAC, and pod security
  • Automate bare-metal provisioning, onboarding, lifecycle, and reclamation
  • Build Terraform providers and modules for GPU cluster infrastructure
  • Define and publish cluster availability, job completion, and provisioning latency SLOs
  • Automate incident management and runbook execution
  • Operate Prometheus, Grafana, Alertmanager, and PagerDuty monitoring
  • Detect GPU node failures and automate drain, cordon, taint, and workload rescheduling

Requirements

  • 5+ years in Kubernetes operations
  • At least 2 years managing GPU workloads on Kubernetes
  • Deep understanding of the Nvidia GPU operator, device plugin, and GPU scheduling
  • Experience with topology-aware scheduling and GPU resource management
  • Experience building multi-tenant Kubernetes platforms with strong isolation
  • Experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom systems
  • Proficiency in Terraform, Helm, and GitOps workflows such as ArgoCD or Flux
  • Strong SRE background in SLI/SLO frameworks, incident management, and capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Strong Go or Python programming skills for operator and CRD development
  • Experience designing or implementing automated Kubernetes remediation
  • Runbook-as-code mindset

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available