Lead Site Reliability Engineer

We are seeking a Lead Site Reliability Engineer to embed with the team and drive reliable operations across a broad infrastructure landscape. You will own monitoring, observability, and logging with Dynatrace and Splunk, and help evolve dashboards, alerting, and analysis practices. You will also join the on-call rotation and partner on incident response. Apply now.

Responsibilities

  • Monitor and sustain the health, performance, and reliability of the client's applications and services
  • Operate and continuously enhance Dynatrace and Splunk, including dashboards, alerting, anomaly detection, and log analysis
  • Analyze alerts, determine likely root causes, and deliver actionable recommendations to engineering and incident management teams
  • Support incident response activities and participate in the on-call rotation
  • Perform Root Cause Analysis (RCA) and drive measurable post-incident improvements
  • Define and monitor SLOs, SLAs, and error budgets
  • Detect and remediate observability and monitoring gaps across services
  • Maintain operational runbooks and supporting documentation

Requirements

  • Proven background with 5+ years in Site Reliability Engineering, Production Operations, DevOps, or a closely related role
  • Deep expertise in Dynatrace and Splunk, including APM, alerting, dashboards, RUM, synthetic monitoring, service flow analysis, SPL queries, and log analysis
  • Hands-on experience with production incident management and on-call support, covering alert triage, incident response, RCA, and post-incident reviews
  • Strong troubleshooting skills and RCA capability across distributed applications and services
  • Solid understanding of application architecture, service dependencies, integrations, performance analysis, dependency mapping, and bottleneck identification
  • Practical experience with AWS services, including CloudWatch, ECS, EC2, ALB, Route53, RDS, and VPC
  • Working knowledge of CI/CD pipelines and release validation processes
  • English proficiency at B2 (Upper-Intermediate) level or higher

Nice to have

  • Familiarity with AI-assisted observability capabilities
  • Experience optimizing monitoring and alerting strategies
  • Exposure to Infrastructure as Code (Terraform or equivalent)
  • Travel/Airline industry experience
  • Background supporting modernization and cloud transformation initiatives

Benefits

  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available