Senior Platform SRE
You own and improve the reliability platform used by engineering teams. You implement observability with OpenTelemetry and distributed tracing, maintain SLOs and error budgets, build deployment and self-healing automation, run chaos experiments, develop CI/CD pipelines, contribute to AI incident tooling, mentor engineers, define SRE standards, guide system design and capacity planning, facilitate post-incident reviews, and track remediation actions.
Responsibilities
- Implement monitoring and observability with OpenTelemetry and distributed tracing
- Maintain SLOs, error budgets, and burn-rate tracking
- Establish 24/7 operational readiness
- Automate deployments and zero-downtime patching
- Engineer auto-remediation and automated traffic rerouting
- Build reliability-focused automation tools and CI/CD pipelines
- Contribute to the SRE AI agent
- Mentor junior SREs and Reliability Champions
- Author and evolve SRE standards
- Guide teams in system design, capacity planning, and architectural reviews
- Design SLOs around customer journeys
- Facilitate blameless post-incident reviews
- Maintain the Lessons Register
- Track remediation actions to closure
- Surface incident patterns quarterly
Requirements
- 6+ years of experience in the stated technical areas
- Hands-on OpenTelemetry experience
- Production experience with Honeycomb, Datadog, Dynatrace, or Grafana
- Ability to instrument Java or Python services
- Experience designing SLIs and setting error budgets
- Experience with multi-window burn-rate alerts
- Experience building CI/CD pipelines with blue/green or canary releases and automated rollback
- Kubernetes experience required
- HashiCorp Nomad experience is advantageous
- Understanding of cloud networking and infrastructure as code
- Production-quality Java or Python coding
- Strong understanding of distributed systems
- On-call production experience
- Experience facilitating blameless post-incident reviews
- Chaos engineering experience
- Experience writing RFCs and building engineering standards
- Experience in high-throughput production environments
- Strong distributed-systems troubleshooting skills
Benefits
- Hybrid working with 3 days in the office
- Tailored development programs
- Mentoring opportunities
- Clear career progression
- Sports and social clubs
- Extra time off for volunteering and community work