SRE L1 Support/Cloud Platform Ops Engineers
You monitor GPU clusters, networks, storage systems, and environmental sensors; respond to alerts and execute incident runbooks; triage and replace hardware; perform standard remediation; collect diagnostics for escalation; manage incident tickets; complete physical data center tasks; conduct shift handoffs; maintain runbooks; and assist with hardware deployment, firmware updates, and inventory management.
Responsibilities
- Monitor GPU cluster health, network status, storage systems, and environmental sensors
- Respond to alerts and execute runbooks for GPU, network, node, and storage incidents
- Identify failed GPUs, NICs, PSUs, disks, and cables
- Execute GPU resets, node drains and reboots, link reseating, and BMC recovery
- Collect logs, DCGM output, network diagnostics, and hardware health reports for escalation
- Manage incident tickets through resolution or escalation
- Perform cable installation, hardware swap-outs, rack and stack, and labeling
- Execute shift handoffs with the APAC operations team
- Maintain and update operational runbooks
- Assist with hardware deployment, firmware updates, and inventory management
Requirements
- 2+ years of experience in NOC, data center operations, or IT support
- Basic Linux system administration
- Familiarity with Prometheus, Grafana, Nagios, or equivalent monitoring tools
- Experience with ServiceNow or Jira Service Management
- Ability to perform rack and stack, cabling, and hardware replacement
- Strong communication skills
- Ability to work 8AM-8PM PST shifts with rotation
Benefits
- Attractive welfare benefits