Cloud Engineer
Key Responsibilities
Incident & SLA Management
- Respond to, triage, and resolve incidents (P1–P4) within contracted SLAs, owning each ticket from detection through to resolution.
- Drive root-cause analysis (RCA) and contribute to post-incident reviews for major incidents.
- Provide clear, timely updates to customers and internal stakeholders throughout the incident lifecycle.
Monitoring, Alerting & Standby
- Serve on the standby / on-call rota and respond promptly to alerts raised by the monitoring platform (Elastic APM and associated observability tooling).
- Proactively monitor dashboards and system health to detect and address issues before they impact service.
- Tune alert thresholds and reduce alert noise to improve signal quality and response efficiency.
Troubleshooting & Cross-TeamCollaboration
- Partner with application, network, security, platform, and other engineering teams to diagnose and resolve complex issues spanning multiple domains.
- Escalate to L3 / specialist teams or vendors (e.g., AWS Support) where required, and manage the escalation through to closure.
Operational Support (BAU)
- Perform day-to-day operational support of AWS environments — compute (EC2), storage (S3 / EBS), networking (VPC, security groups, load balancers), IAM, patching, and backup/restore.
- Execute changes in line with ITIL change management processes and maintain a strong compliance and audit posture.
- Contribute to problem management and continuous service-improvement initiatives.
Documentation & Governance
- Maintain and improve runbooks, standard operating procedures (SOPs), and knowledge-base articles.
- Adhere to Singapore Government security requirements (e.g., IM8) and GCC (Government Commercial Cloud) governance and operating standards.
Required Qualifications & Experience
- Bachelor’s degree in Computer Science, Information Technology, or related field.
- 3–4 years of relevant experience as a Day-2 Cloud / Cloud Operations Engineer supporting production AWS environments.
- Solid working knowledge of core AWS services (EC2, S3, VPC, IAM, CloudWatch, ELB, RDS).
- Hands-on experience with ITIL-aligned incident, problem, and change management in a managed-services or operations context.
- Demonstrated experience supporting Singapore Government accounts.
- GCC (Government Commercial Cloud) experience — familiarity with GCC operating models, controls, and governance.
- Willingness and ability to participate in a standby / on-call and shift rota.
- Strong troubleshooting mindset with the ability to remain composed and effective during high-pressure incidents.
Preferred (Nice-to-Have)
- Experience with Elastic APM / the Elastic Stack (Elasticsearch, Kibana) for application performance monitoring and observability.
- AWS certification — AWS Certified SysOps Administrator – Associate or Solutions Architect – Associate.
- Scripting / automation skills (Python, Bash) and Infrastructure-as-Code (Terraform or CloudFormation).
- ITIL Foundation certification.
- Exposure to other observability / monitoring tools (CloudWatch, Grafana, Prometheus).