> Markdown version of [/jobs/ext/2385800-network-automation-software-lead](https://www.wearedevelopers.com/jobs/ext/2385800-network-automation-software-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Network Automation Software Lead - **Company:** TensorWave Inc. - **Location:** Las Vegas, NV, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Border Gateway Protocol, Code Review, Continuous Integration, Ethernet, White-Box Testing, Python (Programming Language), Netconf, Remote Direct Memory Access, Release Management, Ansible, Prometheus, Software Engineering, Data Logging, Grafana, Backend, Kubernetes, Apache Kafka, Open Network Automation Platform, Software Version Control - **Published:** August 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=99ed3d044407e271 ## About the Role * 8+ years of relevant experience * Proven experience building network automation at scale - ideally at a hyperscaler, large cloud, or large-scale datacenter / AI-infrastructure operator. * You've built or been a core contributor to a ZTP / device-provisioning system end-to-end, not just maintained one. * Strong software engineering fundamentals: Python and/or Go, with real production practices (version control, testing, CI/CD, code review). * Hands-on depth with datacenter Clos fabrics and the protocols that run them: BGP, EVPN/VXLAN, and ideally RoCEv2 / RDMA for GPU networks at scale. * Fluency with modern network automation tech: gNMI/gNOI, OpenConfig/YANG, NETCONF; source-of-truth systems (NetBox / Nautobot); NOS platforms (SONiC/FRR or vendor equivalents); tooling like Nornir / NAPALM / Ansible. * Experience with observability / telemetry pipelines (Prometheus, Grafana, Kafka, OpenTelemetry, or similar). * Comfort running services on Kubernetes / containers. * Leadership: you've led a team or been the clear technical owner of a platform, and you instinctively treat internal users as customers., * Experience with GPU / AI training or inference clusters and their backend networks. * Familiarity with the AMD networking ecosystem (Pensando DPUs, Ultra Ethernet) or building on Ethernet-based RDMA fabrics. * Whitebox / disaggregated networking and SONiC at scale. * Network validation / digital-twin tooling (e.g., Batfish, containerlab). * Multi-site / multi-region datacenter buildouts. ## Description We're hiring a Network Automation Software Lead to build and own the end-to-end zero-touch provisioning (ZTP) and automation platform that stands up and operates our GPU network fabrics and to lead the small team of engineers and SREs building it. The goal is zero human intervention and full fabric validation: a switch goes from rack-and-stack to production-ready automatically, ensuring every link across the fabric is validated against the plan and spec. Our Network Engineering team is your customer; you build the platform, they run the network on top of it. You'll report to a Software Engineering Lead, which means this is real software engineering, not scripts bolted onto a NOC. Source control, testing, CI/CD, code review, release management, and on-call are the baseline. You'll also be responsible for making sure the platform integrates cleanly with the rest of our software and platform tooling. What You'll Do * Own the end-to-end ZTP pipeline: bare-metal switch boot image + base config registration in source of truth full intended config validation production - with zero human intervention. * Build intent-based config generation off a network source of truth / IPAM, with GitOps-style deployment, pre/post-change validation, and safe rollout and rollback. * Establish network validation and pre-deployment testing (snapshot/digital-twin testing) so changes are caught before they hit production fabrics. * Build streaming telemetry and metrics/logging pipelines (gNMI / OpenConfig) for fabric health. * Instrument what matters for GPU networks: RoCE health (PFC/ECN counters), optics and link errors, BGP / EVPN state, capacity and utilization. * Deliver dashboards and alerting the network team actually uses - signal, not noise. * Gather requirements, build self-service APIs and interfaces, and relentlessly accelerate their deployment velocity. * Partner closely so the tooling reflects how the network is actually operated and turned up. * Hire, mentor, and grow a small team of software engineers and SREs; own roadmap, prioritization, and delivery - while still carrying a meaningful share of the code yourself. * Set technical direction and standards, and ensure clean integration points with the broader platform stack (infra provisioning, CI/CD, secrets, identity, existing observability). * Bring software engineering rigor to network automation: code review, testing, release management, and on-call ownership. ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Embracing the Hybrid Cloud: Unlocking Success with Ansible](https://www.wearedevelopers.com/videos/932-embracing-the-hybrid-cloud-unlocking-success-with-ansible) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)