AI Cloud Senior DevOps Engineer
You will operate as the backbone of deployment and infrastructure operations for AI products and platforms. You will automate CI/CD and MLOps workflows, build scalable cloud-native infrastructure, manage GPU resources, implement high availability and disaster recovery, establish observability, enforce security standards, and lead incident resolution.
Responsibilities
- Design and maintain CI/CD pipelines for applications and machine learning models
- Automate build, testing, deployment, and rollback processes
- Build and scale Kubernetes- and Docker-based cloud infrastructure
- Manage GPU clusters and specialized computing resources
- Design high-availability and disaster recovery strategies
- Provision infrastructure with Terraform, Ansible, and Helm
- Build monitoring, logging, and alerting systems
- Establish security, release, secrets management, and compliance standards
- Lead troubleshooting, root cause analysis, and remediation during incidents
Requirements
- Bachelor's degree or above in Computer Science, Engineering, or a related technical field
- 5+ years of experience in DevOps, SRE, or cloud infrastructure roles
- Expert knowledge of Linux and networking principles
- Mastery of Docker and Kubernetes
- Experience with AWS, GCP, Azure, Alibaba Cloud, or other public or hybrid cloud platforms
- Strong coding or scripting skills in Go, Python, Shell, or another major language
- Knowledge of CI/CD, Infrastructure as Code, observability, and SRE
- Experience with MLOps, model serving, GPU clusters, or large-scale distributed systems is preferred