Senior Site Reliability Engineer

You will build reliable, observable, and self-healing infrastructure at scale. You will automate operational tasks and incident response, improve monitoring and alerting, collaborate on resilient system designs, participate in on-call rotations, lead post-incident reviews, and document operational processes and runbooks.

Responsibilities

  • Improve platform reliability and performance
  • Design, build, and maintain tools that automate operational tasks and incident response
  • Implement and improve monitoring, alerting, and tracing solutions
  • Collaborate on scalable and resilient system designs
  • Participate in on-call rotations
  • Lead post-incident reviews
  • Develop and document operational processes and runbooks
  • Contribute to SLO, SLI, and reliability-metric adoption

Requirements

  • Advanced knowledge of Linux/Unix systems in production environments
  • Experience with Kubernetes and container orchestration
  • Proficiency with Terraform and Ansible
  • Experience with Prometheus, Grafana, Loki, or ELK
  • Familiarity with Bash, Python, Go, or Ruby
  • Working knowledge of Git and CI/CD pipelines
  • Understanding of incident management and root cause analysis
  • Knowledge of cloud-native reliability and security best practices

Benefits

  • Paid Time Off
  • Wellhub
  • Annual bonus based on company and team performance
  • Flexible work hours

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available