Network Automation Software Lead

TensorWave Inc.
Las Vegas, NV, United States
29 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Border Gateway Protocol Code Review Continuous Integration Ethernet White-Box Testing Python (Programming Language) Netconf Remote Direct Memory Access Release Management Ansible
+9 more
Prometheus Software Engineering Data Logging Grafana Backend Kubernetes Apache Kafka Open Network Automation Platform Software Version Control

Job description

We’re hiring a Network Automation Software Lead to build and own the end-to-end zero-touch provisioning (ZTP) and automation platform that stands up and operates our GPU network fabrics and to lead the small team of engineers and SREs building it.

The goal is zero human intervention and full fabric validation: a switch goes from rack-and-stack to production-ready automatically, ensuring every link across the fabric is validated against the plan and spec. Our Network Engineering team is your customer; you build the platform, they run the network on top of it.

You’ll report to a Software Engineering Lead, which means this is real software engineering, not scripts bolted onto a NOC. Source control, testing, CI/CD, code review, release management, and on-call are the baseline. You’ll also be responsible for making sure the platform integrates cleanly with the rest of our software and platform tooling.

What You’ll Do

  • Own the end-to-end ZTP pipeline: bare-metal switch boot image + base config registration in source of truth full intended config validation production - with zero human intervention.
  • Build intent-based config generation off a network source of truth / IPAM, with GitOps-style deployment, pre/post-change validation, and safe rollout and rollback.
  • Establish network validation and pre-deployment testing (snapshot/digital-twin testing) so changes are caught before they hit production fabrics.
  • Build streaming telemetry and metrics/logging pipelines (gNMI / OpenConfig) for fabric health.
  • Instrument what matters for GPU networks: RoCE health (PFC/ECN counters), optics and link errors, BGP / EVPN state, capacity and utilization.
  • Deliver dashboards and alerting the network team actually uses - signal, not noise.
  • Gather requirements, build self-service APIs and interfaces, and relentlessly accelerate their deployment velocity.
  • Partner closely so the tooling reflects how the network is actually operated and turned up.
  • Hire, mentor, and grow a small team of software engineers and SREs; own roadmap, prioritization, and delivery - while still carrying a meaningful share of the code yourself.
  • Set technical direction and standards, and ensure clean integration points with the broader platform stack (infra provisioning, CI/CD, secrets, identity, existing observability).
  • Bring software engineering rigor to network automation: code review, testing, release management, and on-call ownership.

Requirements

  • 8+ years of relevant experience

  • Proven experience building network automation at scale - ideally at a hyperscaler, large cloud, or large-scale datacenter / AI-infrastructure operator.
  • You’ve built or been a core contributor to a ZTP / device-provisioning system end-to-end, not just maintained one.
  • Strong software engineering fundamentals: Python and/or Go, with real production practices (version control, testing, CI/CD, code review).
  • Hands-on depth with datacenter Clos fabrics and the protocols that run them: BGP, EVPN/VXLAN, and ideally RoCEv2 / RDMA for GPU networks at scale.
  • Fluency with modern network automation tech: gNMI/gNOI, OpenConfig/YANG, NETCONF; source-of-truth systems (NetBox / Nautobot); NOS platforms (SONiC/FRR or vendor equivalents); tooling like Nornir / NAPALM / Ansible.
  • Experience with observability / telemetry pipelines (Prometheus, Grafana, Kafka, OpenTelemetry, or similar).
  • Comfort running services on Kubernetes / containers.
  • Leadership: you’ve led a team or been the clear technical owner of a platform, and you instinctively treat internal users as customers., * Experience with GPU / AI training or inference clusters and their backend networks.
  • Familiarity with the AMD networking ecosystem (Pensando DPUs, Ultra Ethernet) or building on Ethernet-based RDMA fabrics.
  • Whitebox / disaggregated networking and SONiC at scale.
  • Network validation / digital-twin tooling (e.g., Batfish, containerlab).
  • Multi-site / multi-region datacenter buildouts.

Benefits & conditions

Pulled from the full job description Parental leave 401(k) Health insurance Paid time off Vision insurance Health savings account Dental insurance, * Stock Options

  • 100% paid Medical, Dental, and Vision insurance for Employees
  • Company Health Savings Account Contributions
  • 100% paid Short Term and Long Term Disability Insurance for Employees
  • Life and Voluntary Supplemental Insurance Options
  • Other Insurance Options, such as Pet & Legal Insurance
  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid Holidays
  • Parental Leave
  • Other In-Office Perks

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · World Congress 2024

1:32 min

Grasping core networking and infrastructure automation components

Michael Cade · LIVE

Videos

See all

Related articles

See all