Staff Network Engineer (AI Fabric, Datacenter and Edge Networking)
JobgetherRemotely
bgprocencclrdmagpu clusteringhigh performance networkingnvidia/mellanox
Job Description
📋 Description
- Own long-term technical direction and operational strategy for AI networking.
- Design, deploy, and operate GPU networking fabrics for AI workloads.
- Architect large-scale RoCE and Ethernet fabrics (leaf-spine, fat-tree, rail).
- Lead end-to-end delivery of networking initiatives and drive adoption of standards.
- Build observability and tooling for network reliability and day-two ops.
- Collaborate with compute/platform/SRE teams to shape infrastructure for distributed AI.
🎯 Requirements
- Extensive hands-on experience designing and operating large-scale datacenter networks.
- Expertise in BGP, OSPF, ECMP, EVPN/VXLAN.
- Experience with NVIDIA/Mellanox networking platforms and high-performance interconnects.
- Deep expertise in GPU fabrics for distributed AI workloads and NCCL patterns.
- Hands-on RoCE/RDMA fabric design and tuning for GPU clusters.
- Strong automation skills (Python, Bash) and ability to build reusable tools.
🎁 Benefits
- Attractive compensation package based on experience and impact.
- Flexible and hybrid-friendly working environment.
- Opportunity to work with an internationally diverse team.
- Career growth within a fast-growing technology scale-up.
- Significant autonomy and technical ownership over networking architecture.
- Opportunity to shape networking standards and automation early on.
Back to all jobs