Network Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
The Network Engineering Team is responsible for the design, validation, and ongoing operation of all networking services that underpin both the internal management platform and the customer-facing cloud infrastructure - including high-performance Ethernet fabrics, InfiniBand, WAN connectivity, and DC networking. The team acts as a 3rd/4th line escalation point for the support organisation., As a Senior Network Engineer, you will own the design, automation, and in-service operation of our AI-optimised network fabrics - low-latency, high-bandwidth InfiniBand and Ethernet networks supporting large-scale training and inference workloads. You’ll take ownership of technical areas end to end, act as a senior escalation point, and help raise the bar on how the team automates, documents, and operates the network., * Design, validate, and operate large-scale InfiniBand/RoCE and Ethernet fabric architectures at rack, row, and DC scale, with tight integration to bare-metal provisioning and cluster management systems.
- Apply deep expertise in high-performance Ethernet fabrics (BGP, EVPN, VxLAN, LACP, QoS) and contribute to reference architectures and standards implemented consistently across sites.
- Design and engineer perimeter and security infrastructure - firewalls, NAT, VPN, and security policy architecture - across WAN and DC edge environments.
- Build and maintain network automation in a GitOps model: Python/Ansible tooling for provisioning, configuration validation, and compliance, with version-controlled configuration and CI/CD-driven change across multi-vendor environments.
- Drive operational excellence: resolve complex escalations, lead root-cause analysis for performance and stability issues, and reduce reactive toil through runbooks, automation, and measurable SLOs.
- Improve network observability - telemetry, monitoring, and alerting that give clear visibility into fabric health and traffic patterns.
- Maintain the accuracy of source-of-truth network inventory and configuration data, with all changes flowing through structured change management.
- Collaborate with deployment, DC operations, platform engineering, and vendors on new site delivery, and mentor engineers across the team through reviews and knowledge sharing.
Requirements
- 8+ years of network engineering experience, with depth in HPC, AI, or hyperscale data centre environments.
- Extensive hands-on experience with RDMA-aware networking (InfiniBand, RoCE) for AI/HPC workloads, including subnet managers (OpenSM, UFM) and fabric orchestration.
- Expert-level knowledge of modern DC routing and control planes (BGP, EVPN-VxLAN, Clos/spine-leaf), with production experience on platforms such as Cumulus, Nokia, or Arista EOS.Strong automation skills: Python and Ansible, Git-based workflows, and familiarity with modern IaC and pipeline tooling (Terraform, GitLab CI/GitHub Actions); you treat the network as code rather than managing devices by hand.
- Strong design and engineering experience with firewall platforms - Juniper SRX and/or Palo Alto - including security policy architecture, HA design, and multi-tenant segmentation.
- Experience designing network telemetry and observability for high-throughput environments.
- Proven ability to work cross-functionally with systems, storage, and HPC/AI workload teams, and comfortable leading incident response at senior escalation levels.
- Hands-on, adaptable, and comfortable in a fast-paced environment building next-generation infrastructure for ML scale-out.
Benefits & conditions
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $150,000-$210,000 USD
About the company
Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
At Nscale, our Engineering team plays a critical role in driving the deployment and then subsequent management of our infrastructure and software platforms..
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Everything a Developer Needs to Know About MCP with Neo4j
Highest Paying Tech Companies for Developers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence