SRE L1 Support/Cloud Platform Ops Engineers

You monitor GPU clusters, networks, storage systems, and environmental sensors; respond to alerts and execute incident runbooks; triage and replace hardware; perform standard remediation; collect diagnostics for escalation; manage incident tickets; complete physical data center tasks; conduct shift handoffs; maintain runbooks; and assist with hardware deployment, firmware updates, and inventory management.

Responsibilities

  • Monitor GPU cluster health, network status, storage systems, and environmental sensors
  • Respond to alerts and execute runbooks for GPU, network, node, and storage incidents
  • Identify failed GPUs, NICs, PSUs, disks, and cables
  • Execute GPU resets, node drains and reboots, link reseating, and BMC recovery
  • Collect logs, DCGM output, network diagnostics, and hardware health reports for escalation
  • Manage incident tickets through resolution or escalation
  • Perform cable installation, hardware swap-outs, rack and stack, and labeling
  • Execute shift handoffs with the APAC operations team
  • Maintain and update operational runbooks
  • Assist with hardware deployment, firmware updates, and inventory management

Requirements

  • 2+ years of experience in NOC, data center operations, or IT support
  • Basic Linux system administration
  • Familiarity with Prometheus, Grafana, Nagios, or equivalent monitoring tools
  • Experience with ServiceNow or Jira Service Management
  • Ability to perform rack and stack, cabling, and hardware replacement
  • Strong communication skills
  • Ability to work 8AM-8PM PST shifts with rotation

Benefits

  • Attractive welfare benefits

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available