Senior Site Reliability Engineer (AWS Cloud)

  • Architect & Drive the reliability, scalability, and performance of our multi-cloud provisioning platform across all production stacks.
  • Architect and implement end-to-end automation pipelines to eliminate manual intervention, actively identifying and reducing technical toil.
  • Define, monitor, and improve critical system health indicators (SLIs/SLOs), including latency, throughput, error rates, and capacity usage, making data-driven architectural recommendations.
  • Lead Collaboration with product and cross-functional engineering teams to embed reliability and security considerations early into the software development lifecycle (SDLC).
  • Own Incident Response Management: Design robust detection mechanisms, triage critical incidents, lead deep Root Cause Analysis (RCA), and implement long-term preventative engineering solutions.
  • Simplify Complex Systems: Continually audit platform operations to identify bottlenecks, eliminate single points of failure, and reduce structural complexity.
  • 8+ years of experience working within the cloud environment, in roles such as SRE (Site Reliability Engineer) or Cloud Platform/Reliability Engineer
  • Strong experience in cloud development and multi-cloud environments, preferably with a strong exposure to AWS cloud
  • Knowledge of cloud architecture, scalability, and high-availability design
  • Hands-on experience with Kubernetes and container orchestration
  • Experience with Terraform and Infrastructure as Code (IaC)
  • Experience designing automation and CI/CD pipelines to reduce operational toil
  • Strong understanding of SRE principles, SLIs/SLOs, monitoring, and observability
  • Proven experience with Incident Management, Root Cause Analysis (RCA), and reliability engineering
  • Ability to identify performance bottlenecks, single points of failure, and architectural risks

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available