Senior Front-End Network Engineer, AI Infrastructure Operations

NSCALE, LLC
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Data Analysis Border Gateway Protocol Common Lisp Object Systems Configuration Management Complex Networks Data Centers Extract Transform Load (ETL) Linux Ethernet Firmware Monitoring of Systems
+15 more
Storage Area Network (SAN) Python (Programming Language) Network Architecture Routing Open Shortest Path First (OSPF) Performance Tuning Prometheus Shell Script Data Streaming AI Infrastructure Cloud-native Network Functions (CNF) Computer Network Operations Grafana AI Platforms Front End Software Development

Job description

Within Nscale, the Network Operations team is responsible for the performance and reliability of the high-speed networks that underpin our AI platforms. These front-end networks are critical to inference workloads, cluster management, data movement, and storage connectivity., In this role, you will be responsible for the day-to-day health, stability, and performance of Nscale’s large-scale Ethernet front-end networks. You’ll bring deep operational expertise from hyperscale or high-performance environments and play a key role in incident response, performance tuning, automation, and continuous improvement of production AI networking systems., * Owning the operational health, configuration consistency, and performance tuning of large-scale Ethernet front-end fabrics (leaf-spine / Clos) supporting AI inference, management, and storage workloads

  • Leading the diagnosis and resolution of complex network incidents (P0/P1), spanning optics, routing, switching hardware, long-haul circuits, and storage connectivity layers
  • Driving blameless postmortems and implementing preventative fixes to improve long-term fabric stability and availability
  • Partnering with SREs to define requirements for automation and tooling, and contributing to network provisioning, validation, and monitoring systems
  • Collaborating with Network Architecture and Engineering teams to validate designs and enforce standards for routing, congestion management, firmware baselines, and change safety
  • Monitoring fabric utilisation and performance, identifying bottlenecks, and tuning for predictable latency and throughput on front-end networks
  • Acting as a subject matter expert for cross-functional teams on high-speed Ethernet networking, long-haul/DCI circuits, and storage network integration
  • Participating in an on-call rotation supporting mission-critical, customer-facing infrastructure

Requirements

Do you have experience in Optics?, * 5+ years of experience in network engineering, with at least 3 years operating large-scale Ethernet data centre or cloud networks

  • Deep, hands-on operational experience with high-speed Ethernet fabrics in hyperscale or production environments
  • Strong expertise with Arista (EOS) and/or Nokia (7220 IXR / 7250 IXR / 7750 SR series) platforms
  • Solid understanding of modern data centre networking, including BGP, OSPF, ECMP, EVPN-VXLAN, and leaf-spine architectures
  • Proven experience with long-haul circuits and DCI (dark fiber, carrier Ethernet, coherent optics)
  • Experience with storage networking over Ethernet and shared storage connectivity
  • Proven ability to troubleshoot complex network issues using Linux-based tooling and fabric diagnostics
  • Proficiency in Python, Go, or shell scripting for automation, data analysis, or configuration management
  • Experience working in a 24/7 operational environment with a strong focus on reliability and toil reduction, * Extensive hands-on experience with Arista or Nokia platforms at scale
  • Deep familiarity with front-end network patterns for large AI clusters (inference traffic, management networks, and storage integration)
  • Experience operating large-scale DCI / long-haul optical or carrier networks
  • Strong background in network observability and telemetry systems (streaming telemetry, sFlow, Prometheus, Grafana, etc.)
  • Prior experience in automation-first network operations or building internal tooling

Benefits & conditions

  • Highly competitive package (base + equity) with reviews every 12 months .
  • Join the fastest-growing tech startup , your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
  • Human-First Flexibility: We treat you as humans first . Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life’s moments.

Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work.

About the company

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard ¡ WWC 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet ¡ LIVE

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos ¡ LIVE

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 ¡ WWC Europe 2026

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani ¡ Europe 2026 Virtual

2:42 min

Dissecting artificial intelligence layers from compute to applications

Christian Nagel Christian Nagel +3 ¡ WWC Europe 2026

Videos

See all

Related articles

See all