> Markdown version of [/jobs/ext/2723035-distributed-systems-engineer](https://www.wearedevelopers.com/jobs/ext/2723035-distributed-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Distributed Systems Engineer - **Company:** Institute Inc - **Location:** Sunnyvale, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Software Debugging, Distributed Systems, Fault Tolerance, Github, InfiniBand, PCI Express, Remote Direct Memory Access, System Programming, Graphics Processing Unit (GPU), Computer Network Operations, Pytorch, Large Language Models - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-distributed-systems-engineer-institute-of-foundation-mode-8145987 ## About the Role · Communication-compute overlap and topology-aware collective optimization · Deep debugging of NCCL, RDMA, and custom communication layers · Hybrid expert parallel strategies in modern large-scale MoE systems · Elastic and resilient distributed job orchestration concepts · Congestion analysis and routing optimization across InfiniBand/RoCE fabrics · Microbenchmarking and performance modeling for communication-heavy workloads Expected Technical Depth · Hybrid expert parallel communication for Mixture-of-Experts training · Scaling behavior under network pressure · Distributed orchestration for elastic, large-scale training · Fault detection and recovery in distributed GPU workloads · Cross-layer bottlenecks: GPU * NIC * PCIe * NVSwitch * Fabric * Scheduler Required Background · Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth) · Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA · Deep familiarity with NCCL and/or UCX internals · Strong systems programming ability (C/C++, Rust, or Go) · Strong familiarity with modern model training frameworks such as PyTorch · Ability to troubleshoot and profile training performance issues related to communication bottlenecks · Ability to translate research ideas into production-grade optimizations · Experience debugging distributed hangs, desynchronization, and performance regressions What We Mean by "Hardcore" · You can explain why an communication degrades at scale and how to fix it · You have improved real cluster throughput via communication redesign · You can trace a distributed hang across ranks and identify the root cause · You are comfortable working at the boundary between hardware and runtime, Master's, or Bachelor's + 1 year of relevant experience. ## Description The Institute of Foundation Models (IFM) designs and operates ultra-scale GPU supercomputing systems to train next-generation foundation models. We believe performance, fault tolerance, and scalability are co-designed across model architecture, communication systems, runtime, and hardware topology. This role sits at the core of that effort - driving communication performance, distributed reliability, and cross-layer optimization for large-scale training workloads. The Mission We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including hybrid parallelism and Mixture-of-Experts (MoE) workloads. This is not a network operations role. This is a systems-level engineering position focused on performance engineering, distributed debugging, and communication-runtime co-design. · Design and optimize expert-parallel and hybrid-parallel communication patterns · Drive high-performance hierarchical collectives for MoE workloads · Co-design runtime orchestration with communication topology awareness · Reduce tail latency and improve determinism across thousands of GPUs · Architect fault-tolerant distributed execution under real-world cluster failures, · Include a link to your GitHub (required) · Provide links to relevant distributed systems, HPC, or large-scale training projects · Include a list of publications and/or public technical reports (if applicable) · Describe the hardest distributed debugging problem you solved · Include measurable performance improvements you have delivered ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Data Science & more: The Lopez dilemma](https://www.wearedevelopers.com/magazine/10-data-science-more-the-lopez-dilemma) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)