Senior SRE Engineer

We are seeking a Senior SRE Engineer with strong technical authority to influence design and operational decisions across architecture, engineering, security, and operations teams. The ideal candidate is a pragmatic problem-solver who remains calm and methodical under pressure, balancing reliability, security, cost, and delivery speed while clearly communicating complex technical concepts to diverse audiences. This role requires onsite presence at the client office three times a week.

Responsibilities

  • Design and operate observability platforms, including monitoring, logging, and alerting systems
  • Manage metrics, logs, APM, and alerting through Datadog
  • Apply SRE principles such as SLOs, error budgets, incident management, and reliability engineering
  • Collaborate with security teams to uphold cloud security principles
  • Apply cloud security best practices to ensure system reliability
  • Manage security incidents and drive mitigation and remediation measures
  • Monitor, optimize, and control costs associated with ML models and API-based AI services
  • Operate and maintain Kubernetes and containerized platforms
  • Manage AWS infrastructure, including EKS, ECS, EC2, networking, IAM, and managed services

Requirements

  • 6-9 years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles
  • Strong experience with AWS services such as EKS, ECS, EC2, networking, IAM, and managed services
  • Deep hands-on experience with Kubernetes and containerized platforms
  • Proven experience designing and operating observability platforms, including monitoring, logging, and alerting
  • Hands-on experience with Datadog for metrics, logs, APM, and alerting
  • Strong understanding of SRE principles, including SLOs, error budgets, incident management, and reliability engineering
  • Solid understanding of cloud security principles and experience collaborating with security teams
  • Extensive hands-on experience in cloud security engineering, applying best practices to ensure system reliability, managing security incidents, and driving mitigation and remediation measures
  • Experience or working knowledge of FinOps practices, including monitoring, optimizing, and controlling costs associated with ML models and API-based AI services

Nice to have

  • Experience supporting multi-cloud or hybrid environments
  • Exposure to Infrastructure as Code, such as Terraform and CloudFormation
  • Experience in large-scale, complex, or regulated environments
  • Knowledge of vector databases and RAG architecture for building internal SRE knowledge assistants
  • Knowledge of Generative AI and LLM platforms, such as Claude and Amazon Bedrock

Benefits

Opportunity to work on technical challenges that may impact across geographies

Vast opportunities for self-development: online university, knowledge sharing opportunities globally, learning opportunities through external certifications

Opportunity to share your ideas on international platforms

Sponsored Tech Talks & Hackathons

Unlimited access to LinkedIn learning solutions

Possibility to relocate to any EPAM office for short and long-term projects

Focused individual development

Benefit package:

  • Health benefits
  • Retirement benefits
  • Paid time off
  • Flexible benefits

Forums to explore beyond work passion (CSR, photography, painting, sports, etc.)

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available