> Markdown version of [/jobs/ext/3029114-principal-network-engineer](https://www.wearedevelopers.com/jobs/ext/3029114-principal-network-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Network Engineer - **Company:** Graphcore Limited - **Location:** Austin, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Analysis, Bash Shell, Border Gateway Protocol, Common Lisp Object Systems, Computer Clusters, Configuration Management, Computer Engineering, Network Congestion, Data Centers, Distributed Computing Environment, Ethernet, InfiniBand, Python (Programming Language), Network Architecture, Open Shortest Path First (OSPF), Remote Direct Memory Access, Data Streaming, AI Infrastructure, High Performance Computing, Computer Network Technologies, Nx-os, Data Center Networking, Infrastructure Automation Frameworks, Information Technology, Low Latency, Cisco, Golang - **Published:** September 22, 2026 - **Apply:** https://job-boards.greenhouse.io/graphcore/jobs/8463815002 ## About the Role * BS or MS or equivalent experience in Computer Science, Electrical Engineering, Network Engineering, or related technical discipline. * 12+ years of progressive network engineering experience with at least 3 years in hyperscale, high-density, or HPC data center environments. * Expert-level knowledge of data center routing and switching protocols including BGP, OSPF, and EVPN-VXLAN architectures. * Strong operational understanding of RDMA networking technologies such as RoCEv2 or InfiniBand. * Hands-on experience with modern merchant silicon networking platforms and NOS platforms such as Arista EOS, Cisco NX-OS, or SONiC. * Experience deploying high-speed network technologies including 400G/800G optics and large-scale fabric architectures. * Proficiency in automation and scripting languages such as Python, Go, Bash, or similar tools. * Strong collaboration and communication skills across cross-functional engineering teams. Desirable * Experience operating large-scale AI or GPU clusters. * Familiarity with network telemetry frameworks and streaming analytics. * Experience implementing NetDevOps workflows and infrastructure automation pipelines. * Experience influencing vendor roadmaps or evaluating next-generation networking technologies. ## Description We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next-generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero-packet-loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high-performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore's large-scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high-speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high-performance network fabrics that power Graphcore's AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and inference workloads. Engineers work on pioneering technologies including high-speed Ethernet fabrics, lossless networking, RDMA transport, and large-scale automation frameworks to support next-generation AI clusters., * Assist in defining ultra-high-bandwidth, non-blocking AI network fabrics (Clos spine-leaf-super-spine architectures) for large-scale distributed AI workloads. * Optimize performance of lossless Ethernet fabrics using congestion control mechanisms such as PFC, ECN, and DCQCN to support RDMA/RoCEv2 communication. * Lead initiatives to implement NetDevOps practices and develop automation for provisioning, configuration management, and network remediation. * Design and deploy high-resolution telemetry pipelines to monitor network health, detect microbursts, and analyze congestion patterns. * Support modeling, deployment, configuration, and monitoring of data center network fabrics including scale-out, scale-up, and front-end networks. * Collaborate cross-functionally with hardware engineers, AI researchers, and data center operations teams to co-design high-performance infrastructure. * Provide technical leadership and mentorship to network engineers while establishing best practices and operational standards. * Contribute to the long-term networking strategy and roadmap for Graphcore's AI infrastructure. * Research and evaluate next-generation high-speed networking technologies and vendor solutions. ## Related Videos - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Computer Vision from the Edge to the Cloud done easy](https://www.wearedevelopers.com/videos/263-computer-vision-from-the-edge-to-the-cloud-done-easy) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)