Site Reliability Engineer (SRE)
Roles & Responsibilities
Ensure high availability, reliability, scalability, and performance of production applications and infrastructure.
Define and monitor SLIs, SLOs, and SLAs for critical services.
Participate in on-call rotations, incident response, troubleshooting, and root-cause analysis.
Identify operational risks and implement automation to reduce manual intervention and recurring incidents.
Conduct post-incident reviews and track corrective/preventive actions.
Design, implement, and maintain CI/CD pipelines for application and infrastructure deployments.
Integrate Git-based source control with automated build, test, security scanning, and deployment workflows.
Skills/Requirements
Strong experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.
Strong knowledge of Linux, shell scripting, Git, CI/CD, Docker, and Kubernetes.
Experience with observability platforms, logging, metrics, tracing, alerting, and dashboarding.
Experience integrating systems with enterprise monitoring, alerting, SIEM, or Incident Response workflows.
Experience defining and implementing runbooks, operational procedures, escalation paths, and production support models.