> Markdown version of [/jobs/ext/1431910-hpc-network-engineer](https://www.wearedevelopers.com/jobs/ext/1431910-hpc-network-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC Network Engineer - **Company:** Fuse Limited - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** IEEE 802.1X, Artificial Intelligence, Border Gateway Protocol, Network Operating System (NOS), Common Lisp Object Systems, Configuration Management, Computer Networks, Network Congestion, Data Centers, Software Debugging, Ethernet, Network Interface Controllers, InfiniBand, Storage Area Network (SAN), Virtual Private Networks (VPN), Python (Programming Language), Linux System Administration, Network Planning and Design, Operational Databases, Performance Tuning, Remote Direct Memory Access, Remote Access Technology, Broadcom, Ansible, Prometheus, Data Streaming, Encapsulation (Networking), Wireless Networks, Datadog, Dynamic Routing, Grafana, Bare Metal, Fortinet, Open Network Automation Platform, Cisco Switches, Nvme - **Published:** July 25, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=ba8c3de2537ca42b ## About the Role * 5+ years as a network engineer operating production data centre networks * Strong dynamic routing experience (BGP in particular), plus overlay/encapsulation design and troubleshooting (e.g. EVPN/VXLAN) * Hands-on experience with leaf-spine / Clos fabric design and operation * Experience with modern data centre network operating systems and comfortable in the Linux networking stack, not just a vendor CLI * Practical RDMA fabric experience: lossless Ethernet (e.g. RoCEv2 with PFC/ECN/DCQCN tuning) or InfiniBand, with an understanding of why lossless behaviour matters for GPU workloads * Network automation as a working practice, not an aspiration: scripting (e.g. Python), configuration management (e.g. Ansible), config generation from a source of truth, version-controlled changes * Solid Linux administration fundamentals: you can debug from the host side as well as the switch side * Experience with network telemetry and monitoring (e.g. Prometheus/Grafana, sFlow/IPFIX, streaming telemetry) * Experience running corporate/campus networks: wired and wireless, switching, NAC/802.1X, VPN and remote access (e.g. Cisco Catalyst/Meraki or comparable) * Clear communicator who enjoys teaching: able to document, pair, and run sessions that bring less network-savvy colleagues up to speed Nice to have * Experience with GPU cluster networking specifically (e.g. NVIDIA Spectrum-X or Quantum InfiniBand, ConnectX/BlueField NICs and DPUs, UFM, SHARP, or equivalent Broadcom/Arista AI fabric platforms) * Container networking experience: CNI plugins and BGP integration between clusters and the fabric. * Understanding of collective-communication libraries and how fabric behaviour shows up as training/inference performance * Multi-tenant network design: VRF-based isolation, tenant bandwidth guarantees, secure shared infrastructure * Experience with enterprise firewall platforms (e.g. FortiGate, Palo Alto), including HA deployment and virtualised/segmented instances * Storage networking experience (e.g. NVMe-oF, lossless storage fabrics, per-tenant storage isolation) * Bare-metal provisioning environments (e.g. MAAS, PXE, Redfish) and how network bootstrap fits into node lifecycle * Optical layer knowledge at 200/400/800G: transceivers, MPO cabling, link qualification * Experience standing up a data centre network from greenfield * Relevant certifications (e.g. CCNP/CCIE or equivalent), valued as evidence of depth, not a gate ## Description * Design and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control, and buffer tuning at scale * Build and manage leaf-spine data centre fabrics, with routed underlay and overlay design (e.g. BGP, EVPN/VXLAN) * Implement and maintain per-tenant network isolation across compute, storage, and management planes. * Automate network provisioning, configuration, and validation, treating switch config as code (e.g. Ansible, Python, NetBox as source of truth), deployed through CI * Build telemetry and observability for the fabric: flow-level and buffer-level visibility, dashboards, and alerting that catches congestion and link degradation before tenants do (e.g. Prometheus/Grafana/Datadog, streaming telemetry) * Troubleshoot performance issues end to end, from optics and cabling through switch buffers to NIC/DPU configuration and collective-communication behaviour on the hosts * Operate the out-of-band management network, console access, and remote recovery paths * Support tenant onboarding: segmentation and addressing, bandwidth and isolation guarantees, and capacity planning as the cluster scales * Write clear design documentation capturing decisions, rationale, and rejected alternatives * Own and maintain the office network: wired and wireless infrastructure, firewalling, VPN/remote access, and connectivity between the office and data centre environments * Upskill colleagues on networking: share knowledge through documentation, run-throughs, and pairing so the wider team can operate and troubleshoot the fabric confidently ## Related Videos - [The OpenTelemetry mistakes I keep seeing (and how to stop making them)](https://www.wearedevelopers.com/videos/100158-the-opentelemetry-mistakes-i-keep-seeing-and-how-to-stop-making-them) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Demystifying application networking in the cloud](https://www.wearedevelopers.com/videos/675-demystifying-application-networking-in-the-cloud) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023) - [What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?](https://www.wearedevelopers.com/magazine/701-what-makes-wearedevelopers-world-congress-different-from-every-other-tech-event)