> Markdown version of [/jobs/ext/575758-sr-staff-software-engineer-hpc-network](https://www.wearedevelopers.com/jobs/ext/575758-sr-staff-software-engineer-hpc-network). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Staff Software Engineer - HPC Network... - **Company:** LinkedIn Corporation - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $181,000.0 - $297,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, C++ (Programming Language), Common Lisp Object Systems, Computer Clusters, Profiling, Computer Networks, Network Congestion, Linux, Distributed Systems, Ethernet, Network Interface Controllers, Python (Programming Language), Machine Learning, Network Architecture, Network Protocols, Performance Tuning, Remote Direct Memory Access, Systems Architecture, Data Processing, High Performance Computing, Large Language Models, Backend, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Low Latency, Apache Flink, Apache Kafka, Spark Streaming, Stream Processing, Data Pipelines, Golang, Programming Languages, Microservices - **Published:** June 21, 2026 - **Apply:** https://www.juju.com/job/00000000g9nwpx ## About the Role + BA/BS Degree in Computer Science or related technical discipline, or equivalent practical experience + 10+ years of experience building and operating large-scale distributed systems or data-intensive backend platforms. + Experience in one or more programming languages such as Go, Python, C++, or similar. + Experience in Linux system engineering and host networking. + Demonstrated knowledge of network protocols, fabric design, and performance optimization. + Proven ability to lead complex technical initiatives end-to-end in a multi-team environment. + Experience with system design skills with focus on scalability, reliability, and performance. + Experience with container platforms (Kubernetes) and microservices. Preferred Qualifications: + Experience supporting large-scale AI or HPC workloads. + Familiarity with LLM training frameworks and communication libraries (e.g., NCCL, MPI). + Experience with streaming systems (Kafka, Flink, Spark Streaming, or similar) and high-throughput data pipeline architectures. + Experience with performance benchmarking and profiling tools. + Experience with infrastructure automation or configuration management tools. + Demonstrated influence across organizations (tech lead, architect, principal/IC leadership roles). Suggested Skills: + Distributed Systems + HPC Networking + Performance Optimization + Technical Leadership ## Description We are seeking an HPC Network Engineer to design, deploy, and operate high-performance, low-latency Ethernet fabrics for large-scale GPU clusters. The role focuses on RoCE v2-based GPU interconnect networks supporting AI/ML training, inference, and HPC workloads. You will work closely with systems, GPU, platform, and software teams to build scalable, lossless Ethernet networks optimized for RDMA traffic. As a Senior Staff Software Engineer, you will define long-term technical direction, lead cross-org initiatives, mentor senior engineers, and drive solutions for complex distributed systems challenges at massive scale. This role requires deep expertise in backend systems, data processing, and large-scale system design, with strong understanding of networking concepts. Responsibilities: + Network architecture and design for large-scale LLM training and inference workloads. + Design RoCE v2-based GPU interconnection fabrics for multi-rack and multi-pod GPU clusters + Define lossless Ethernet architectures (Clos / fat-tree / leaf-spine) optimized for RDMA + Select and validate 400G / 800G Ethernet switching platforms and NICs (ConnectX, BlueField, etc.) + Deep expertise in host-level and Kubernetes pod networking architectures, including enablement of high-performance features such as RDMA and GPU Direct. + Experience in host network performance tuning for large-scale collective communications, balancing latency, throughput, and congestion control. + Analyze system performance and diagnose complex cross-layer issues., A request for an accommodation will be responded to within three business days. However, non-disability related requests, such as following up on an application, will not receive a response. LinkedIn will not discharge or in any other manner discriminate against employees or applicants because they have inquired about, discussed, or disclosed their own pay or the pay of another employee or applicant. However, employees who have access to the compensation information of other employees or applicants as a part of their essential job functions cannot disclose the pay of other employees or applicants to individuals who do not otherwise have access to compensation information, unless the disclosure is (a) in response to a formal complaint or charge, (b) in furtherance of an investigation, proceeding, hearing, or action, including an investigation conducted by LinkedIn, or (c) consistent with LinkedIn's legal duty to furnish information. San Francisco Fair Chance Ordinance Pursuant to the San Francisco Fair Chance Ordinance, LinkedIn will consider for employment qualified applicants with arrest and conviction records. Pay Transparency Policy Statement As a federal contractor, LinkedIn follows the Pay Transparency and non-discrimination provisions described at this link: https://lnkd.in/paytransparency. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Super scaling for the Super Bowl: How to survive 30 million users hitting your backend in 30 minutes](https://www.wearedevelopers.com/videos/100356-super-scaling-for-the-super-bowl-how-to-survive-30-million-users-hitting-your-backend-in-30-minutes) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)