GPU DC East-West Network SRE Expert (SME)
You operate InfiniBand and RoCEv2 fabrics carrying GPU communication across large data centers. You manage fabric topologies, subnet management, congestion control, NCCL tuning, firmware, telemetry, fault diagnosis, and vendor escalations while turning network incidents and mitigations into predictive models and automated workflows.
Responsibilities
- Operate InfiniBand fat-tree, rail-optimized, and dragonfly fabrics
- Operate RoCEv2 networks across Nvidia, Arista, and Cisco platforms
- Manage IB monitoring, diagnostics, and subnet management with UFM
- Monitor and tune IB and RoCE performance
- Configure adaptive routing, DCQCN, ECN, and traffic isolation
- Tune NCCL topology detection, ring and tree algorithms, and GDR
- Manage switch and HCA firmware lifecycles
- Diagnose link flaps, symbol errors, packet drops, routing anomalies, and credit stalls
- Coordinate vendor support escalations, bug reports, and RMAs
- Integrate IB/RoCE telemetry into the collection pipeline
- Define link and straggler predictors
- Convert incidents and mitigations into labeled examples and automated workflows
Requirements
- 5+ years in data center networking
- At least 3 years focused on InfiniBand or RoCE fabrics
- Experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
- Understanding of IB subnet management, partitioning, and QoS
- Experience deploying RoCEv2 with PFC, ECN, and DCQCN
- Proficiency with UFM or equivalent IB fabric management tools
- Knowledge of 400G/800G optics, cabling standards, and structured cabling
- Experience diagnosing IB/RoCE issues with ibdiagnet, perfquery, and ibstat
- Understanding of NCCL and GPU communication topology
- Experience building dashboards or alerts on RDMA counters, or ability to define fabric-health model features
- Runbook-as-code mindset