SRE L1 Support/Cloud Platform Ops Engineer
You provide front-line monitoring and incident response for GPU data centers during an 8AM–8PM PST shift. You execute runbooks, triage hardware and infrastructure issues, collect diagnostics, manage tickets, perform physical data center tasks, coordinate handoffs, update procedures, and supply structured incident data for automation.
Responsibilities
- Monitor GPU cluster health, network status, storage systems, and environmental sensors
- Respond to alerts and execute runbooks for common incidents
- Triage failed GPUs, NICs, PSUs, disks, and cables
- Perform GPU resets, node drains and reboots, link reseating, and BMC recovery
- Collect logs, DCGM output, network diagnostics, and hardware health reports
- Manage incident tickets through resolution or escalation
- Install cables, swap hardware, rack and stack equipment, and label components
- Perform structured shift handoffs with the APAC operations team
- Maintain and update operational runbooks
- Assist with hardware deployment, firmware updates, and inventory management
- Tag and document novel incidents for automation
- Provide structured handoff notes
Requirements
- 2+ years in a NOC, data center operations, or IT support role
- Basic Linux system administration
- Familiarity with Prometheus, Grafana, Nagios, or equivalent monitoring tools
- Experience with ServiceNow or Jira Service Management
- Ability to perform rack and stack, cabling, and hardware replacement
- Strong communication skills for handoffs, incident documentation, and escalation
- Ability to work the 8AM–8PM PST shift schedule with 12-hour shifts and rotation
- Curiosity about automation
- Comfort with structured data and incident ticketing