GPU DC East-West Network SRE Expert (SME)

You operate InfiniBand and RoCEv2 fabrics carrying GPU communication across large data centers. You manage fabric topologies, subnet management, congestion control, NCCL tuning, firmware, telemetry, fault diagnosis, and vendor escalations while turning network incidents and mitigations into predictive models and automated workflows.

Responsibilities

  • Operate InfiniBand fat-tree, rail-optimized, and dragonfly fabrics
  • Operate RoCEv2 networks across Nvidia, Arista, and Cisco platforms
  • Manage IB monitoring, diagnostics, and subnet management with UFM
  • Monitor and tune IB and RoCE performance
  • Configure adaptive routing, DCQCN, ECN, and traffic isolation
  • Tune NCCL topology detection, ring and tree algorithms, and GDR
  • Manage switch and HCA firmware lifecycles
  • Diagnose link flaps, symbol errors, packet drops, routing anomalies, and credit stalls
  • Coordinate vendor support escalations, bug reports, and RMAs
  • Integrate IB/RoCE telemetry into the collection pipeline
  • Define link and straggler predictors
  • Convert incidents and mitigations into labeled examples and automated workflows

Requirements

  • 5+ years in data center networking
  • At least 3 years focused on InfiniBand or RoCE fabrics
  • Experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
  • Understanding of IB subnet management, partitioning, and QoS
  • Experience deploying RoCEv2 with PFC, ECN, and DCQCN
  • Proficiency with UFM or equivalent IB fabric management tools
  • Knowledge of 400G/800G optics, cabling standards, and structured cabling
  • Experience diagnosing IB/RoCE issues with ibdiagnet, perfquery, and ibstat
  • Understanding of NCCL and GPU communication topology
  • Experience building dashboards or alerts on RDMA counters, or ability to define fabric-health model features
  • Runbook-as-code mindset

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available