Principal Network Engineer

Nscale Ltd.
London, UK
8 days ago
Apply on uk.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Link Aggregation (Ethernet) Artificial Intelligence Border Gateway Protocol Big Data Common Lisp Object Systems Cloud Computing Continuous Integration Data Centers Network Address Translation Ethernet Github InfiniBand
+20 more
Subnetting Virtual Private Networks (VPN) Python (Programming Language) Network Security Network Configuration and Change Management Network Architecture Network Service Overlay Transport Virtualization Remote Direct Memory Access Ansible AI Infrastructure Computer Networking Systems High Performance Computing Juniper AI Platforms Gitlab-ci Git Flow Low Latency Terraform Open Network Automation Platform

Job description

We are hiring a Principal Network Engineer to act as a senior technical authority for Nscale’s AI-optimised network infrastructure.

Our Network Engineering team is responsible for the design, validation, and ongoing operation of the networking services underpinning both our internal management platform and customer-facing cloud infrastructure. This includes high-performance Ethernet fabrics, InfiniBand, RoCE, WAN connectivity, and large-scale data centre networking.

In this role, you’ll set technical direction across the low-latency, high-bandwidth networks supporting large-scale AI training and inference workloads. You’ll own critical technical domains end-to-end and help raise the bar for architecture, automation, operational rigour, and engineering standards across Nscale.

This is a deeply technical Principal-level role combining hands-on engineering with broad architectural influence. You’ll define reference architectures, drive consistency across sites, lead complex technical decisions and escalations, and mentor engineers while partnering closely with Deployment, Data Centre Operations, Platform Engineering, Systems, Storage, and technology vendors., * Define, design, validate, and evolve large-scale InfiniBand, RoCE, and Ethernet fabric architectures at rack, row, and data centre scale.

  • Design networks that integrate closely with bare-metal provisioning and cluster management systems.
  • Own technical direction for high-performance Ethernet fabrics, including BGP, EVPN, VXLAN, LACP, and QoS.
  • Establish reference architectures and engineering standards that can be implemented consistently across Nscale’s data centre estate.
  • Identify systemic risks and architectural gaps and drive durable solutions that improve scalability, reliability, and operational simplicity., * Lead Nscale’s network automation strategy using a GitOps operating model.
  • Build and guide Python and Ansible tooling for provisioning, configuration validation, compliance, and operational workflows.
  • Drive version-controlled configuration and CI/CD-based network change across multi-vendor environments.
  • Apply Infrastructure-as-Code and Network-as-Code principles to reduce manual intervention and improve operational consistency.
  • Continuously identify opportunities to automate repetitive operational tasks and reduce reactive toil., * Design and engineer perimeter and network security infrastructure across WAN and data centre edge environments.
  • Own architecture across firewalls, NAT, VPN, security policies, and multi-tenant segmentation.
  • Design highly available and scalable security architectures appropriate for mission-critical AI infrastructure., * Lead complex technical escalations and root-cause analysis for network performance, reliability, and stability issues.
  • Establish measurable SLOs and operational standards for network services.
  • Set technical direction for network observability, telemetry, monitoring, and alerting.
  • Ensure clear visibility into fabric health, traffic patterns, performance, and capacity.
  • Develop runbooks, automation, and engineering improvements that systematically reduce operational toil.
  • Act as a senior 3rd/4th line escalation point for complex networking issues., * Ensure the accuracy and reliability of source-of-truth network inventory and configuration data.
  • Establish structured engineering and change-management practices for network configuration.
  • Ensure network changes are controlled, auditable, repeatable, and scalable across multiple sites., * Partner with Deployment, Data Centre Operations, Platform Engineering, Systems, Storage, and vendors on new site delivery and platform evolution.
  • Lead architecture and design reviews for significant network initiatives.
  • Mentor engineers and raise technical capability across the wider networking organisation.
  • Lead complex technical decisions and incidents spanning networking, systems, storage, and AI/HPC workloads.
  • Influence engineering strategy and standards across teams without relying on formal authority., Expect a dynamic progression plan tailored to your ambitions. Shape network architecture, establish engineering standards, solve complex infrastructure challenges, and influence the technical direction of Nscale’s global AI platform.

Requirements

  • 10+ years of network engineering experience, with significant depth in HPC, AI, hyperscale, or large-scale data centre environments.
  • Extensive hands-on experience with RDMA-aware networking for AI/HPC workloads, including InfiniBand and/or RoCE.
  • Experience with subnet managers and fabric orchestration technologies such as OpenSM or NVIDIA UFM.
  • Expert-level understanding of modern data centre routing and control planes, including BGP, EVPN-VXLAN, and Clos/spine-leaf architectures.
  • Production experience with network platforms such as Cumulus, Nokia, or Arista EOS.
  • Strong network automation expertise using Python and Ansible.
  • Experience with Git-based workflows and modern Infrastructure-as-Code and CI/CD tooling such as Terraform, GitLab CI, or GitHub Actions.
  • Deep experience designing and engineering firewall infrastructure using platforms such as Juniper SRX and/or Palo Alto.
  • Experience designing telemetry and observability solutions for high-throughput, performance-sensitive environments.
  • Proven ability to lead complex technical decisions and incidents across multiple engineering disciplines.

Architecture & Technical Leadership

  • Demonstrated experience defining network architecture, engineering standards, and technical strategy beyond a single project or data centre.
  • Ability to balance performance, reliability, operability, scalability, and delivery velocity when making technical decisions.
  • Experience influencing engineering teams and technical direction without formal authority.
  • Strong communication skills with the ability to explain complex technical trade-offs to engineers, operators, and senior stakeholders.
  • Proven ability to mentor and develop other engineers through design reviews, technical guidance, incident leadership, and knowledge sharing., * Deeply technical and comfortable remaining hands-on at Principal level.
  • Highly analytical with a structured approach to solving complex infrastructure problems.
  • Strong sense of ownership and accountability.
  • Comfortable operating in a fast-paced environment where priorities and requirements evolve quickly.
  • Pragmatic and able to balance architectural excellence with business and delivery requirements.
  • Passionate about building next-generation infrastructure for AI and ML at scale.

Benefits & conditions

At Nscale, you’ll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We’re building something extraordinary, and we want you at the core.

Highly competitive package (base + equity) with reviews every 12 months.

Join one of the fastest-growing AI infrastructure companies - your opportunity to design and build the high-performance networks powering some of the world’s most demanding AI workloads.

About the company

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

At Nscale, our Engineering team plays a critical role in designing, deploying, and operating the infrastructure and software platforms that power our customers and enable AI workloads at scale.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on uk.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

1:23 min

Closing thoughts and educational resources for edge network engineering

Austin Gil · LIVE

Videos

See all

Related articles

See all