Network Engineer
You will design, deploy, and operate the network infrastructure supporting GPU AI factories across Europe. You will manage physical cabling, high-speed Ethernet and InfiniBand fabrics, routing, overlays, transport tuning, automation pipelines, and production troubleshooting for large GPU clusters and live AI training workloads.
Responsibilities
- Design and deploy InfiniBand NDR 400G, HDR, and high-speed Ethernet fabrics for GPU clusters of 1,000+ nodes
- Configure and operate Arista, Juniper, and Mellanox/NVIDIA equipment
- Manage BGP, OSPF, and VXLAN overlays
- Tune RoCE and InfiniBand transport for NCCL and UCX workloads
- Maintain network automation pipelines across all sites
- Troubleshoot performance regressions, packet loss, and congestion during live AI training runs
Requirements
- 4+ years of datacenter networking experience
- Experience with InfiniBand or 400G/800G Ethernet at scale
- Deep familiarity with RDMA, RoCE v2, and GPU training cluster communication patterns
- Solid knowledge of Linux networking internals including DSCP, ECN, PFC, and adaptive routing
- Experience with Ansible, Terraform, Netbox, and CI/CD pipelines
- Ability to interpret tcpdump, perftest, and ib_write_bw diagnostics and correlate them with application performance