Network Engineer

NuScale Power Corporation
United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$150,000.0 - $210,000.0
Working hours
Regular working hours
Job source

Tech stack

Link Aggregation (Ethernet) Artificial Intelligence Border Gateway Protocol Common Lisp Object Systems Cloud Computing Cyber Security Continuous Integration Data Centers Network Address Translation Ethernet Github InfiniBand
+14 more
Subnetting Virtual Private Networks (VPN) Python (Programming Language) Overlay Transport Virtualization Remote Direct Memory Access Ansible Wide Area Networks High Performance Computing Juniper Gitlab-ci Git Flow Low Latency Terraform Open Network Automation Platform

Job description

The Network Engineering Team is responsible for the design, validation, and ongoing operation of all networking services that underpin both the internal management platform and the customer-facing cloud infrastructure - including high-performance Ethernet fabrics, InfiniBand, WAN connectivity, and DC networking. The team acts as a 3rd/4th line escalation point for the support organisation., As a Senior Network Engineer, you will own the design, automation, and in-service operation of our AI-optimised network fabrics - low-latency, high-bandwidth InfiniBand and Ethernet networks supporting large-scale training and inference workloads. You’ll take ownership of technical areas end to end, act as a senior escalation point, and help raise the bar on how the team automates, documents, and operates the network., * Design, validate, and operate large-scale InfiniBand/RoCE and Ethernet fabric architectures at rack, row, and DC scale, with tight integration to bare-metal provisioning and cluster management systems.

  • Apply deep expertise in high-performance Ethernet fabrics (BGP, EVPN, VxLAN, LACP, QoS) and contribute to reference architectures and standards implemented consistently across sites.
  • Design and engineer perimeter and security infrastructure - firewalls, NAT, VPN, and security policy architecture - across WAN and DC edge environments.
  • Build and maintain network automation in a GitOps model: Python/Ansible tooling for provisioning, configuration validation, and compliance, with version-controlled configuration and CI/CD-driven change across multi-vendor environments.
  • Drive operational excellence: resolve complex escalations, lead root-cause analysis for performance and stability issues, and reduce reactive toil through runbooks, automation, and measurable SLOs.
  • Improve network observability - telemetry, monitoring, and alerting that give clear visibility into fabric health and traffic patterns.
  • Maintain the accuracy of source-of-truth network inventory and configuration data, with all changes flowing through structured change management.
  • Collaborate with deployment, DC operations, platform engineering, and vendors on new site delivery, and mentor engineers across the team through reviews and knowledge sharing.

Requirements

  • 8+ years of network engineering experience, with depth in HPC, AI, or hyperscale data centre environments.
  • Extensive hands-on experience with RDMA-aware networking (InfiniBand, RoCE) for AI/HPC workloads, including subnet managers (OpenSM, UFM) and fabric orchestration.
  • Expert-level knowledge of modern DC routing and control planes (BGP, EVPN-VxLAN, Clos/spine-leaf), with production experience on platforms such as Cumulus, Nokia, or Arista EOS.Strong automation skills: Python and Ansible, Git-based workflows, and familiarity with modern IaC and pipeline tooling (Terraform, GitLab CI/GitHub Actions); you treat the network as code rather than managing devices by hand.
  • Strong design and engineering experience with firewall platforms - Juniper SRX and/or Palo Alto - including security policy architecture, HA design, and multi-tenant segmentation.
  • Experience designing network telemetry and observability for high-throughput environments.
  • Proven ability to work cross-functionally with systems, storage, and HPC/AI workload teams, and comfortable leading incident response at senior escalation levels.
  • Hands-on, adaptable, and comfortable in a fast-paced environment building next-generation infrastructure for ML scale-out.

Benefits & conditions

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $150,000-$210,000 USD

About the company

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

At Nscale, our Engineering team plays a critical role in driving the deployment and then subsequent management of our infrastructure and software platforms..

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:23 min

Closing thoughts and educational resources for edge network engineering

Austin Gil · LIVE

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

2:51 min

Designing hardware infrastructure and networking for distributed compute

Anshul Jindal Anshul Jindal +1 · World Congress 2025

Videos

See all

Related articles

See all