Senior Site Reliability Engineer

You will build and scale internal platform offerings covering compute, storage, and networking services to ensure reliability and performance. You will design and implement monitoring, alerting, and incident response systems, collaborate with application software engineers to guide scalable designs, and act as an agent of change to incrementally improve systems as the platform expands globally. You will work with a stack of Python, Java, Terraform, gRPC, Docker, Kubernetes, and Postgres running on AWS, and you will use AI tools daily to reduce toil.

Responsibilities

  • Build and scale internal platform offerings including compute, storage, and networking services
  • Design and implement monitoring, alerting, and incident response systems
  • Collaborate with application software engineers to guide scalable design
  • Drive incremental improvements to systems as the company expands globally

Requirements

  • Extensive experience with cloud platforms such as AWS, Google Cloud Platform, or Azure, including EC2, S3, RDS, and Lambda
  • Experience with Kubernetes or other container orchestration preferred
  • Proficiency with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation
  • Experience with networking concepts including Container Network Interface (CNI) and network policy implementations
  • Experience with proxies and service mesh is a plus
  • Strong knowledge of monitoring tools such as Prometheus, Grafana, ELK Stack, or Datadog
  • Proficiency in Python with the ability to write efficient, maintainable, and scalable code
  • Experience designing, deploying, and maintaining API services with RESTful and/or GraphQL design principles
  • Fluency with AI tools, including building agents to reduce toil
  • Experience operating CI/CD and its best practices appreciated

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available