Infrastructure and MLOps Engineer
You will develop, own, and maintain tools and services that support AI research and engineering teams. You will deploy and maintain services with Kubernetes and Docker, manage cloud infrastructure with Terraform, and improve build, testing, deployment, and productisation processes for machine learning software components.
Responsibilities
- Develop, own, and maintain tools and services for AI research and engineering teams
- Deploy and maintain services with Kubernetes and Docker
- Manage cloud infrastructure with Terraform
Requirements
- Knowledge of Python
- Familiarity with cloud services such as AWS
- Experience managing or developing in Linux environments
- Understanding of CI/CD principles
- Experience using Kubernetes
- Experience maintaining machine learning applications, deploying ML orchestration tools, or managing ML accelerator hardware
- Experience with Infrastructure as Code tools such as Terraform or OpenTofu
- Experience with GitHub Actions
- Experience with modern observability tooling such as Prometheus
- Experience with Grafana
- Knowledge of Go, Java, C++, or a similar language
Benefits
- Flexible working
- Generous annual leave policy
- Private medical insurance
- Health cash plan
- Dental plan
- Pension matched up to 5%
- Life assurance
- Income protection
- Generous parental leave policy
- Employee assistance programme
- Healthy food and snacks
- On-site barista bar