Engineering Team Leader (Site Reliability Engineering)

We are building XTB – a global investment company offering innovative technological solutions that allow our clients to effectively manage their finances in multiple ways. All of this within a single, intuitive XTB app already used by over one million users worldwide!

We are a certified Great Place to Work company.

We are looking for an Engineering Team Leader to drive the development and growth of the Site Reliability Engineering team. In this role, you will have the opportunity to shape the technical and operational direction of SRE practices, lead a resilience strategy, and play a key role in defining and delivering solutions that ensure the reliability and scalability of XTB systems for millions of clients across a growing organization.

  • Professional Background: Several years of experience in SRE, Infrastructure, or DevOps roles managing high-scale, distributed environments.
  • Leadership Experience: Proven track record in a formal management role, leading, mentoring, and developing high-performing SRE or DevOps engineering teams.
  • Infrastructure & Reliability: Extensive experience building and maintaining scalable, reliable, and observable infrastructure systems (Azure, Kubernetes, On-prem).
  • Reliability Strategy: Demonstrated ability to deliver end-to-end reliability strategies, drive architectural improvements, and manage large-scale technical projects from design to production.
  • Cross-Functional Collaboration: Proven ability to drive cultural change, act as a strategic partner to product engineering teams, and work effectively with distributed/remote teams.


Leadership & Management

  • Execution: Break down complex projects into actionable tasks; drive incremental value.
  • Growth: Mentor and support the professional growth of team members.
  • Collaboration: Facilitate workshops; build technical community; resolve conflicts effectively.
  • Operational Excellence: Drive operational excellence and reliability culture within the team; lead incident management and champion post-mortems.
  • Strategy: Proactively manage technical debt; align team output with organizational goals.


Technical skills we expect

  • Programming & Scripting: Strong Python skills for building scalable automation, internal tools, and scripts.
  • Cloud & Orchestration: Expertise in managing Kubernetes, configuration management with Ansible, and designing resilient infrastructure on Azure and on-prem.
  • Observability Engineering: Deep proficiency in building standardized telemetry systems. Mastery of tools like Prometheus, Grafana, OTEL, ELK, Tempo, Thanos, and similar.
  • AI & Automation: Using AI/ML for AIOps, anomaly detection, log analysis, and optimizing reliability workflows.


Nice to have

  • Experience with commercial APM platforms (e.g., Datadog, Splunk, New Relic) and chaos engineering tooling.
  • Proficiency with cloud cost management and FinOps principles.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available