> Markdown version of [/jobs/ext/2007037-principal-ai-network-hardware-systems-engineer](https://www.wearedevelopers.com/jobs/ext/2007037-principal-ai-network-hardware-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal AI Network Hardware Systems Engineer - **Company:** Microsoft - **Location:** Hillsboro, OR, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Automation of Tests, Microsoft Azure, Data Link, Communications Protocols, Computer Engineering, Network Congestion, Software Debugging, Distributed Computing Environment, Network Interface Controllers, Firmware, Networking Hardware, Interoperability, Routing, Network Service, Remote Direct Memory Access, System Software, TCP/IP, AI Infrastructure, Diagnostic Tools, Computer Networking Systems, AI Platforms, Low Latency - **Published:** August 9, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/principal-ai-network-hardware-systems-engineer-hillsboro-or-usa-58865992 ## About the Role and networking validation strategies covering functionality, performance, scale, interoperability, resiliency, and reliability * Characterize network behavior under AI workloads * Evaluate latency, bandwidth utilization, congestion events, and workload communication patterns * Create and automate network stress, scale, and performance qualification methodologies * Lead end-to-end troubleshooting across physical, data link, network, and transport layers * Perform packet-level analysis and protocol debugging using telemetry and diagnostic tools * Investigate switch, NIC, RDMA, routing, congestion control, and protocol issues * Drive corrective actions and long-term reliability improvements using fleet telemetry and lab validation * Build and improve network observability, diagnostics, telemetry, and monitoring solutions * Develop tools for network validation, performance analysis, and failure detection * Improve engineering productivity through automated testing, qualification, and network a years assessment frameworks Tasks * 8+ years of NW HW development * 8+ years of GPU based SU/SO development * 8+ years hands on experience with HS interface architecture and development * Master's Degree in Electrical Engineering * or Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field with 8+ years of experience Key requirements * ## Description Experteer Overview In this role you lead the architecture and deployment of networking infrastructure for Microsoft's MAIA AI platform, enabling high-performance AI training and inference at hyperscale. You will influence architectural direction across silicon, firmware, hardware, and Azure teams, delivering scalable networking solutions from concept to datacenter rollout. You will work on high-speed networking, AI fabric technologies, and distributed training workloads to improve latency, throughput, and reliability. This is a chance to shape next-generation AI networking for a global cloud leader with a strong culture of collaboration and growth. Compensation / Benefits * Define and develop networking requirements for large-scale AI training and inference clusters * Collaborate across silicon, system software, firmware, hardware, and Azure teams to deliver scalable networking solutions from concept through datacenter deployment * Participate in architecture reviews and influence next-generation AI networking roadmaps * Define concepts of operation, telemetry requirements, and operational models for AI infrastructure * Lead design and validation of IP-based AI networking solutions spanning TCP/IP, UDP, routing, QoS, and traffic engineering * Analyze transport-layer behavior and performance across large-scale distributed AI workloads * Evaluate network protocol implementations and debug issues impacting latency, throughput, scalability, and reliability * Drive optimization of network communication paths supporting distributed AI training and inference * Design, validate, and optimize RDMA-based networking solutions for AI clusters * Analyze RDMA performance, congestion behavior, packet loss, retransmissions, and collective communication efficiency * Work with networking vendors and software teams to optimize AI fabric performance and workload scalability * Develop validation methodologies for AI traffic patterns and collective communication workloads * Develop and execute networking validation strategies covering functionality, performance, scale, interoperability, resiliency, and reliability * Characterize network behavior under AI workloads * Evaluate latency, bandwidth utilization, congestion events, and workload communication patterns * Create and automate network stress, scale, and performance qualification methodologies * Lead end-to-end troubleshooting across physical, data link, network, and transport layers * Perform packet-level analysis and protocol debugging using telemetry and diagnostic tools * Investigate switch, NIC, RDMA, routing, congestion control, and protocol issues * Drive corrective actions and long-term reliability improvements using fleet telemetry and lab validation * Build and improve network observability, diagnostics, telemetry, and monitoring solutions * Develop tools for network validation, performance analysis, and failure detection * Improve engineering productivity through automated testing, qualification, and network health assessment frameworks Tasks * 8+ years of NW HW development * 8+ years of GPU based SU/SO development * 8+ years hands on experience with HS interface architecture and development * Master's Degree in Electrical Engineering * or Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field with 8+ years of experience Key requirements * ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)