Senior Platform SRE

You own and improve the reliability platform used by engineering teams. You implement observability with OpenTelemetry and distributed tracing, maintain SLOs and error budgets, build deployment and self-healing automation, run chaos experiments, develop CI/CD pipelines, contribute to AI incident tooling, mentor engineers, define SRE standards, guide system design and capacity planning, facilitate post-incident reviews, and track remediation actions.

Responsibilities

  • Implement monitoring and observability with OpenTelemetry and distributed tracing
  • Maintain SLOs, error budgets, and burn-rate tracking
  • Establish 24/7 operational readiness
  • Automate deployments and zero-downtime patching
  • Engineer auto-remediation and automated traffic rerouting
  • Build reliability-focused automation tools and CI/CD pipelines
  • Contribute to the SRE AI agent
  • Mentor junior SREs and Reliability Champions
  • Author and evolve SRE standards
  • Guide teams in system design, capacity planning, and architectural reviews
  • Design SLOs around customer journeys
  • Facilitate blameless post-incident reviews
  • Maintain the Lessons Register
  • Track remediation actions to closure
  • Surface incident patterns quarterly

Requirements

  • 6+ years of experience in the stated technical areas
  • Hands-on OpenTelemetry experience
  • Production experience with Honeycomb, Datadog, Dynatrace, or Grafana
  • Ability to instrument Java or Python services
  • Experience designing SLIs and setting error budgets
  • Experience with multi-window burn-rate alerts
  • Experience building CI/CD pipelines with blue/green or canary releases and automated rollback
  • Kubernetes experience required
  • HashiCorp Nomad experience is advantageous
  • Understanding of cloud networking and infrastructure as code
  • Production-quality Java or Python coding
  • Strong understanding of distributed systems
  • On-call production experience
  • Experience facilitating blameless post-incident reviews
  • Chaos engineering experience
  • Experience writing RFCs and building engineering standards
  • Experience in high-throughput production environments
  • Strong distributed-systems troubleshooting skills

Benefits

  • Hybrid working with 3 days in the office
  • Tailored development programs
  • Mentoring opportunities
  • Clear career progression
  • Sports and social clubs
  • Extra time off for volunteering and community work

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available