> Markdown version of [/jobs/ext/2669724-staff-network-engineer-ai-fabric-datacenter-and-edge-networking-radian-arc](https://www.wearedevelopers.com/jobs/ext/2669724-staff-network-engineer-ai-fabric-datacenter-and-edge-networking-radian-arc). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Network Engineer (AI Fabric, Datacenter and Edge Networking) - Radian Arc - **Company:** Jobgether - **Location:** Germany (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Border Gateway Protocol, Configuration Management, Complex Networks, Computer Networks, Network Congestion, Wavelength-Division Multiplexing, Data Centers, DDoS Mitigation, Linux, Distributed Computing Environment, Ethernet, Firmware, Network Topologies, Networking Hardware, Python (Programming Language), Open Shortest Path First (OSPF), Operational Data Store, Performance Tuning, Software Architecture, Remote Direct Memory Access, Software Engineering, Virtual Local Area Networks, Infrastructure Automation Frameworks, Low Latency, Operational Systems - **Published:** September 1, 2026 - **Apply:** https://www.adzuna.de/details/5861243381 ## About the Role * Network engineering experience: Strong hands-on experience designing and operating large-scale datacenter networks in production, with expert knowledge of BGP, OSPF, ECMP, and EVPN/VXLAN. * AI networking expertise: Deep experience designing and operating high-performance GPU networking fabrics for distributed AI workloads, including practical RoCE/RDMA expertise. * GPU communication knowledge: Strong understanding of NCCL communication patterns-including all-reduce, all-gather, broadcast, and reduce-scatter-and their implications for network topology and performance. * RoCE optimization: Experience tuning large-scale RoCE fabrics, including PFC and ECN congestion management, and diagnosing NCCL stalls, RDMA congestion, fabric hotspots, and packet-loss issues. * Production infrastructure: Proven experience with high-speed Ethernet and NVIDIA/Mellanox networking platforms, as well as high-performance interconnects, optics, and networking hardware. * Systems troubleshooting: Ability to investigate complex cross-layer problems spanning hardware, firmware, Linux/kernel networking, and distributed application communication. * Automation: Strong Python and/or Bash skills, with experience applying software engineering practices to infrastructure automation and building reusable operational tooling. * Observability: Experience designing network observability and independently analyzing operational data, ideally using tools such as PostHog or equivalent systems. * Architecture and execution: Demonstrated ability to own both long-term architecture and hands-on implementation in lean or rapidly scaling environments. * Technical leadership: Proven ability to lead complex initiatives, establish engineering standards, influence multiple teams without formal authority, and balance performance, reliability, scalability, operability, and cost. * Communication: Strong stakeholder management and communication skills, with the ability to explain architectural decisions, trade-offs, and technical risks clearly. * Mentoring: Ability to raise the technical bar of adjacent engineering teams and contribute to the development of a future networking organization. * Languages: Professional working proficiency in English; the ability to collaborate effectively across international teams is essential. ## Description This is a high-impact Staff-level networking role responsible for the architecture, deployment, and operation of infrastructure powering GPU-based AI workloads. You will own the technical direction of high-performance AI fabrics while remaining deeply hands-on across implementation, troubleshooting, automation, and operations. The scope spans RoCE/RDMA GPU fabrics, datacenter routing, edge connectivity, security, and inter-datacenter backbone infrastructure. As the organization's first dedicated networking specialist, you will establish reusable architectures, engineering standards, and operational practices from the ground up. You will work closely with compute, storage, platform, SRE, observability, and datacenter teams to make networking a core part of the infrastructure design. The role combines architectural leadership with direct execution in a lean, fast-scaling, international environment. It is an opportunity to shape the networking foundation for distributed AI training and inference across global datacenter and edge deployments. Accountabilities * AI fabric architecture: Design, deploy, and operate high-performance GPU networking fabrics for distributed AI workloads, including RoCE, RDMA, Spectrum-X, leaf-spine, fat-tree, rail, and multi-plane architectures. * AI performance optimization: Analyze GPU communication patterns and optimize east-west traffic, congestion control, latency, throughput, and reliability for distributed training and inference workloads. * Datacenter networking: Own Layer 2/3 architecture using BGP, ECMP, EVPN/VXLAN, VLAN/VRF, OVS/OVN, Linux networking, bridges, overlays, and scalable routing patterns. * Security and edge connectivity: Design and operate north-south connectivity, WAF and application-layer protections, TLS termination, DDoS mitigation, and gateway infrastructure. * Inter-datacenter networking: Design private connectivity, dark-fiber and metro-fiber rings, high-capacity WAN links, DWDM transport, inter-site BGP routing, redundancy, and failure-domain isolation. * End-to-end delivery: Lead infrastructure initiatives from architecture and lab validation through BOM validation, datacenter layouts, deployment, acceptance, and production rollout. * Operational excellence: Establish automation for provisioning, configuration management, monitoring, lifecycle management, validation, and day-2 operations while improving reliability and recovery practices. * Incident leadership: Act as the senior escalation point for complex network incidents, leading investigation, root-cause analysis, remediation, and long-term architectural improvements. * Performance and reliability: Define and monitor SLAs, SLOs, latency, reliability, capacity, and operational metrics, translating findings into concrete engineering improvements. * Cross-functional leadership: Partner with infrastructure, platform, SRE, compute, storage, observability, and datacenter teams while setting networking standards and influencing long-term architecture. * Technical leadership: Serve as the primary networking design authority, mentor engineers in adjacent domains, and establish reusable patterns that can scale with the organization. * Automation and leverage: Build Python and/or Bash tooling, validation frameworks, and operational systems that reduce manual work and increase consistency across deployments. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Agent Smith Gets Hardware: Autonomous IoT Hacking From Debug Port to Cloud API](https://www.wearedevelopers.com/videos/100258-agent-smith-gets-hardware-autonomous-iot-hacking-from-debug-port-to-cloud-api) ## Related Articles - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)