HPC Network Engineer

Fuse Limited
Greater London, UK
17 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

IEEE 802.1X Border Gateway Protocol Network Operating System (NOS) Common Lisp Object Systems Configuration Management Computer Networks Network Congestion Data Centers Software Debugging Ethernet Network Interface Controllers InfiniBand
+18 more
Storage Area Network (SAN) Virtual Private Networks (VPN) Python (Programming Language) Linux System Administration Performance Tuning Remote Direct Memory Access Remote Access Technology Ansible Prometheus Data Streaming Encapsulation (Networking) Wireless Networks Dynamic Routing Grafana Bare Metal Fortinet Open Network Automation Platform Nvme

Job description

  • Design and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control and buffer tuning at scale
  • Build and manage leaf-spine data centre fabrics, with routed underlay and overlay design (e.g. BGP, EVPN/VXLAN)
  • Implement and maintain per-tenant network isolation across compute, storage and management planes
  • Automate network provisioning, configuration and validation, treating switch config as code (e.g. Ansible, Python, NetBox as source of truth), deployed through CI
  • Build telemetry and observability for the fabric: flow-level and buffer-level visibility, dashboards and alerting that catch congestion and link degradation before tenants do
  • Troubleshoot performance issues end to end, from optics and cabling through switch buffers to NIC/DPU configuration and collective-communication behaviour on the hosts
  • Operate the out-of-band management network, console access and remote recovery paths
  • Support tenant onboarding: segmentation and addressing, bandwidth and isolation guarantees, and capacity planning as the cluster scales
  • Write clear design documentation capturing decisions, rationale and rejected alternatives
  • Own and maintain the office network: wired and wireless infrastructure, firewalling, VPN/remote access and office-to-data-centre connectivity
  • Upskill colleagues on networking through documentation, run-throughs and pairing

Requirements

  • Strong dynamic routing experience (BGP in particular), plus overlay/encapsulation design and troubleshooting (e.g. EVPN/VXLAN)
  • Hands-on experience with leaf-spine / Clos fabric design and operation
  • Experience with modern data centre network operating systems and comfortable in the Linux networking stack, not just a vendor CLI
  • Practical RDMA fabric experience: lossless Ethernet (e.g. RoCEv2 with PFC/ECN/DCQCN tuning) or InfiniBand, understanding why lossless behaviour matters for GPU workloads
  • Network automation as a working practice: scripting (e.g. Python), configuration management (e.g. Ansible), config generation from a source of truth, version-controlled changes
  • Solid Linux administration fundamentals: you can debug from the host side as well as the switch side
  • Experience with network telemetry and monitoring (e.g. Prometheus/Grafana, sFlow/IPFIX, streaming telemetry)
  • Experience running corporate/campus networks: wired and wireless, switching, NAC/802.1X, VPN and remote access
  • Clear communicator who enjoys teaching
  • Bonus: GPU cluster networking (NVIDIA Spectrum-X or Quantum InfiniBand, ConnectX/BlueField NICs and DPUs, UFM, SHARP or equivalent); container networking (CNI, BGP integration); collective-communication libraries; multi-tenant VRF-based isolation; enterprise firewalls (FortiGate, Palo Alto); storage networking (NVMe-oF); bare-metal provisioning (MAAS, PXE, Redfish); optical layer at 200/400/800G; greenfield data centre network build; CCNP/CCIE or equivalent

Benefits & conditions

  • Competitive salary and eligibility for equity
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Private health insurance
  • Breakfast and dinner allowance for office-based employees

About the company

Fuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.

We’ve raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.

We’re building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.

You’ll design, deploy and operate the network fabric for our multi-tenant AI cluster: the high-performance compute and storage fabrics carrying RDMA traffic between GPUs, the tenant-facing and management networks, firewalling and tenant isolation, and the out-of-band infrastructure that keeps it all recoverable. Beyond the data centre, you’ll own the office network and act as the networking authority for the company. You own the fabric from architecture through day-2 operations.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:23 min

Closing thoughts and educational resources for edge network engineering

Austin Gil · LIVE

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

2:43 min

Bursting GPU capacity over hybrid Kubernetes networks

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

Videos

See all

Related articles

See all