> Markdown version of [/jobs/ext/2499582-senior-network-engineer-gpu-cluster-networking](https://www.wearedevelopers.com/jobs/ext/2499582-senior-network-engineer-gpu-cluster-networking). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Network Engineer - GPU Cluster Networking - **Company:** Advanced Micro Devices, Inc. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $173,600.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Application Performance Management, Border Gateway Protocol, Common Lisp Object Systems, Cloud Computing, Computer Clusters, Computer Networks, Computer Engineering, System Configuration, Data Centers, Microprocessors, Distributed Systems, Ethernet, Network Interface Controllers, Data-Flow Analysis, Monitoring of Systems, Network Topologies, Subnetting, Junos, Automation of Marketing, Network Architecture, Network Planning and Design, Routing, Network Segmentation, Operational Databases, PCI Express, Performance Tuning, Queue Management Systems, Remote Direct Memory Access, Prometheus, Simple Network Management Protocols, Software Deployment, Data Streaming, Virtual Local Area Networks, Network Routing, Graphics Processing Unit (GPU), High Performance Computing, Performance Testing, Computer Network Technologies, Grafana, Backend, Juniper, Data Center Networking, Kubernetes, Information Technology, Low Latency, Data Analytics, Slurm - **Published:** August 29, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88163471/1 ## About the Role You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,. You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions. You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations., * Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments. * Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments. * Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation * Strong hands-on experience with RDMA and RoCEv2 in production environments. * Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet. * Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures. * Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN. * Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality. * Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms. * Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting * Experience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments. * Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies. * Strong hands-on experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting * Experience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads. * Experience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures., * Bachelor's or Master's degree in Computer Engineering, or a related field, or equivalent practical experience. ## Description We are seeking a Senior Network Engineer to join the AMD IT System Engineering team. This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance. The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments. The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems. You will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance., * Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters. * Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs. * Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric. * Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies. * Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures. * Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting. * Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN. Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack. Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks. Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations, AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [A Technical Introduction to Bitcoin's 2nd Layer- The Lightning Network](https://www.wearedevelopers.com/videos/15-a-technical-introduction-to-bitcoin-s-2nd-layer-the-lightning-network) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)