Senior Site Reliability Engineer (SRE Team)

We’re the SRE Team, the specialists behind the reliability of Semrush's robust infrastructure and applications. Our team is a collaborative partner that works with cross-functional teams to identify potential points of failure and weaknesses across the platform. We propose and implement solutions to enhance the reliability of critical systems, making them more resilient to outages and faults.

We believe in the power of a resilient infrastructure, and we're growing our team to achieve even greater things. If you're an engineer excited about identifying points of failure, implementing reliability solutions, and working at the intersection of infrastructure and application resilience, we want to meet you.

Key Responsibilities:

  • Lead the changes in common engineering practices in the Company.

  • Induce application failures and work to recover them from that state.

  • Debug applications using metrics and add traces/metrics as needed.

  • Establish and refine SLOs, cost dashboards, and security hardening initiatives in partnership with stakeholders to guarantee service reliability and performance.

  • Collaborate with development teams to design and implement scalable, reliable, and efficient system architecture.

  • Designs full-stack platform solutions from concept to production.

  • Builds sophisticated tooling in Go/Python to automate operations.

  • Mentors engineers, interviews candidates, leads critical incidents.

  • On-call rotation: Typically one week every 2–3 weeks and may include overnight incidents.

About you

Move together. Raise the bar. Learn fast—grow faster. That’s the default. And here’s what else is needed to succeed in this role:

  • 3+ years of experience as a Site Reliability Engineer.

  • Experience with Kubernetes and Cloud providers.

  • Experience with engineering in Python or Go.

  • Strong understanding of what an application failure is and how to handle it.

  • Ability to debug applications using metrics.

  • Familiarity with traces, observability, and implementation quirks in code.

  • Willingness to be on call and work flexible hours.

  • Team player with good communication abilities.

Not required, but a plus

  • GCP knowledge.

About the perks

  • Unlimited PTO

  • Hobby & team building budget allowance

  • Employee Support Program

  • Loss of family member financial aid

  • Employee Resource Groups

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available