Site Reliability Engineer (SRE)

Role Summary

The Site Reliability Engineer (SRE) at iCareManager (iCM) is responsible for ensuring the reliability, scalability, performance, and availability of production systems. The role blends software engineering and operations, with a strong focus on automation, observability, incident management, and proactive reliability engineering.

SREs enable engineering teams to deliver features rapidly without compromising system stability, especially in regulated healthcare environments.

Key Responsibilities

1. Reliability Engineering

  • Define, implement, and enforce Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets

  • Design and review resilience patterns (redundancy, failover, graceful degradation)

  • Perform capacity planning, load modeling, and scalability analysis

  • Conduct chaos testing and failure injection to identify system weaknesses

  • Reduce Mean Time to Recovery (MTTR) through architectural improvements and tooling

2. Observability & Monitoring

  • Instrument systems with metrics, logs, and distributed traces

  • Build and maintain dashboards that reflect system health and performance

  • Design alerting strategies that are actionable and minimize alert fatigue

  • Identify leading indicators of failure before customer impact


3. Incident Management & Postmortems

  • Participate in and lead production incident response

  • Coordinate with engineering and infrastructure teams during incidents

  • Lead blameless postmortems and document root cause analysis

  • Track and remediate reliability debt and systemic risks



Requirements

Required Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or equivalent experience

  • 3+ years experience in SRE, DevOps, Platform Engineering, or similar roles

  • Strong programming experience in one or more languages (e.g., Dot net and or node JS

  • Hands-on experience with Linux-based systems

  • Experience with cloud platforms (Azure preferred; AWS/GCP acceptable)

  • Solid understanding of networking, distributed systems, and system design

  • Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK, datadog)

Preferred Qualifications

  • Experience in healthcare or regulated environments

  • Familiarity with containerization and orchestration (Docker, Kubernetes)

  • Experience with CI/CD pipelines and infrastructure as code

  • Understanding of security best practices in production systems

  • Experience supporting SOC2-compliant environments

Key Competencies

  • Strong problem-solving and analytical skills

  • Calm and effective during high-pressure incidents

  • Excellent documentation and communication skills

  • Ownership mindset and bias toward automation

  • Collaborative and proactive approach



Benefits

  • A dynamic and collaborative work environment.
  • Opportunities for professional growth and skill development.
  • Competitive salary and benefits package.
  • The chance to play a key role in revolutionizing the healthcare technology industry.


See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available