Sr Ethernet AI Network Engineer

Job Cloud Inc.
United States
21 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Border Gateway Protocol Big Data Network Operating System (NOS) Command-Line Interface Complex Networks Network Congestion Data Centers Linux Ethernet Firmware Transport Layer
+8 more
Network Architecture Routing Packet Analyzer Reliability Engineering Prometheus Virtual Local Area Networks Grafana Open Network Automation Platform

Job description

Sr Ethernet AI Network Engineer to provide senior network engineering services for Ethernet based AI data center fabrics supporting GPU compute clusters, storage connectivity, and management plane services. This role requires strong operational judgment, advanced troubleshooting capabilities, and the ability to resolve complex network issues within large scale production environments. This is a 12 month remote contract opportunity within the United States.

Responsibilities:

  • Deploy, validate, and support Ethernet fabrics used for AI and HPC cluster environments.
  • Troubleshoot Layer 1 through Layer 4 issues involving optics, transceivers, cabling, link bring up, VLANs, MLAG, ECMP, BGP, underlay and overlay reachability, congestion, and packet loss.
  • Validate network readiness for distributed training workloads and large east west traffic patterns.
  • Diagnose performance issues related to buffer pressure, microbursts, PFC behavior, Quality of Service policy, MTU mismatches, routing instability, and oversubscription.
  • Partner with Linux, storage, and cluster deployment teams to isolate host versus network fault domains.
  • Review and execute change plans for switch provisioning, firmware upgrades, topology expansion, and maintenance activities.
  • Capture packet level and counter based evidence to support root cause analysis.
  • Develop operational standards for cable mapping, port policy consistency, and fabric health validation.

Requirements

  • 7 or more years of experience in large scale data center networking, including high bandwidth Ethernet fabrics.
  • Strong experience with spine leaf architectures, routing, switching, and production troubleshooting.
  • Hands on experience with BGP, EVPN, VXLAN, MLAG, ECMP, Quality of Service, Priority Flow Control, RoCE considerations, and telemetry analysis.
  • Experience validating optics, breakout configurations, cable plant integrity, and port level consistency.
  • Proven ability to troubleshoot distributed application impacts caused by network behavior.
  • Experience using switch command line interfaces, automation tools, and packet and counter analysis workflows.
  • Strong documentation skills for topology diagrams, incident timelines, and remediation planning., * Direct experience supporting AI fabrics carrying large scale GPU collective traffic.
  • Familiarity with SONiC, Cumulus Linux, or similar network operating systems used in AI data centers.
  • Experience with streaming telemetry, Prometheus, Grafana, and network site reliability engineering operating models.

Tools and Technologies:

  • Ethernet Data Center Fabrics
  • Spine Leaf Network Architectures
  • BGP
  • EVPN
  • VXLAN
  • MLAG
  • ECMP
  • Quality of Service
  • Priority Flow Control
  • RoCE
  • SONiC
  • Cumulus Linux
  • Prometheus
  • Grafana
  • Network Telemetry Platforms
  • Linux
  • Packet Analysis Tools
  • Network Automation Tooling

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:23 min

Closing thoughts and educational resources for edge network engineering

Austin Gil · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all