> Markdown version of [/jobs/ext/2526945-sr-ethernet-ai-network-engineer](https://www.wearedevelopers.com/jobs/ext/2526945-sr-ethernet-ai-network-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr Ethernet AI Network Engineer - **Company:** Job Cloud Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Border Gateway Protocol, Big Data, Network Operating System (NOS), Command-Line Interface, Complex Networks, Network Congestion, Data Centers, Linux, Ethernet, Firmware, Transport Layer, Network Architecture, Routing, Packet Analyzer, Reliability Engineering, Prometheus, Virtual Local Area Networks, Grafana, Open Network Automation Platform - **Published:** August 15, 2026 - **Apply:** https://www.dice.com/job-detail/98300c74-f030-41de-91c2-f98686292402 ## About the Role * 7 or more years of experience in large scale data center networking, including high bandwidth Ethernet fabrics. * Strong experience with spine leaf architectures, routing, switching, and production troubleshooting. * Hands on experience with BGP, EVPN, VXLAN, MLAG, ECMP, Quality of Service, Priority Flow Control, RoCE considerations, and telemetry analysis. * Experience validating optics, breakout configurations, cable plant integrity, and port level consistency. * Proven ability to troubleshoot distributed application impacts caused by network behavior. * Experience using switch command line interfaces, automation tools, and packet and counter analysis workflows. * Strong documentation skills for topology diagrams, incident timelines, and remediation planning., * Direct experience supporting AI fabrics carrying large scale GPU collective traffic. * Familiarity with SONiC, Cumulus Linux, or similar network operating systems used in AI data centers. * Experience with streaming telemetry, Prometheus, Grafana, and network site reliability engineering operating models. Tools and Technologies: * Ethernet Data Center Fabrics * Spine Leaf Network Architectures * BGP * EVPN * VXLAN * MLAG * ECMP * Quality of Service * Priority Flow Control * RoCE * SONiC * Cumulus Linux * Prometheus * Grafana * Network Telemetry Platforms * Linux * Packet Analysis Tools * Network Automation Tooling ## Description Sr Ethernet AI Network Engineer to provide senior network engineering services for Ethernet based AI data center fabrics supporting GPU compute clusters, storage connectivity, and management plane services. This role requires strong operational judgment, advanced troubleshooting capabilities, and the ability to resolve complex network issues within large scale production environments. This is a 12 month remote contract opportunity within the United States. Responsibilities: * Deploy, validate, and support Ethernet fabrics used for AI and HPC cluster environments. * Troubleshoot Layer 1 through Layer 4 issues involving optics, transceivers, cabling, link bring up, VLANs, MLAG, ECMP, BGP, underlay and overlay reachability, congestion, and packet loss. * Validate network readiness for distributed training workloads and large east west traffic patterns. * Diagnose performance issues related to buffer pressure, microbursts, PFC behavior, Quality of Service policy, MTU mismatches, routing instability, and oversubscription. * Partner with Linux, storage, and cluster deployment teams to isolate host versus network fault domains. * Review and execute change plans for switch provisioning, firmware upgrades, topology expansion, and maintenance activities. * Capture packet level and counter based evidence to support root cause analysis. * Develop operational standards for cable mapping, port policy consistency, and fabric health validation. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [A Technical Introduction to Bitcoin's 2nd Layer- The Lightning Network](https://www.wearedevelopers.com/videos/15-a-technical-introduction-to-bitcoin-s-2nd-layer-the-lightning-network) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)