Distributed Systems Engineer

Institute Inc
Sunnyvale, United States
2 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

C++ (Programming Language) Software Debugging Distributed Systems Fault Tolerance Github InfiniBand PCI Express Remote Direct Memory Access System Programming Graphics Processing Unit (GPU) Computer Network Operations Pytorch
+1 more
Large Language Models

Job description

The Institute of Foundation Models (IFM) designs and operates ultra-scale GPU supercomputing systems to train next-generation foundation models. We believe performance, fault tolerance, and scalability are co-designed across model architecture, communication systems, runtime, and hardware topology. This role sits at the core of that effort - driving communication performance, distributed reliability, and cross-layer optimization for large-scale training workloads. The Mission We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including hybrid parallelism and Mixture-of-Experts (MoE) workloads. This is not a network operations role. This is a systems-level engineering position focused on performance engineering, distributed debugging, and communication-runtime co-design. ยท Design and optimize expert-parallel and hybrid-parallel communication patterns ยท Drive high-performance hierarchical collectives for MoE workloads ยท Co-design runtime orchestration with communication topology awareness ยท Reduce tail latency and improve determinism across thousands of GPUs ยท Architect fault-tolerant distributed execution under real-world cluster failures, ยท Include a link to your GitHub (required) ยท Provide links to relevant distributed systems, HPC, or large-scale training projects ยท Include a list of publications and/or public technical reports (if applicable) ยท Describe the hardest distributed debugging problem you solved ยท Include measurable performance improvements you have delivered

Requirements

ยท Communication-compute overlap and topology-aware collective optimization ยท Deep debugging of NCCL, RDMA, and custom communication layers ยท Hybrid expert parallel strategies in modern large-scale MoE systems ยท Elastic and resilient distributed job orchestration concepts ยท Congestion analysis and routing optimization across InfiniBand/RoCE fabrics ยท Microbenchmarking and performance modeling for communication-heavy workloads Expected Technical Depth ยท Hybrid expert parallel communication for Mixture-of-Experts training ยท Scaling behavior under network pressure ยท Distributed orchestration for elastic, large-scale training ยท Fault detection and recovery in distributed GPU workloads ยท Cross-layer bottlenecks: GPU * NIC * PCIe * NVSwitch * Fabric * Scheduler Required Background ยท Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth) ยท Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA ยท Deep familiarity with NCCL and/or UCX internals ยท Strong systems programming ability (C/C++, Rust, or Go) ยท Strong familiarity with modern model training frameworks such as PyTorch ยท Ability to troubleshoot and profile training performance issues related to communication bottlenecks ยท Ability to translate research ideas into production-grade optimizations ยท Experience debugging distributed hangs, desynchronization, and performance regressions What We Mean by โ€œHardcoreโ€ ยท You can explain why an communication degrades at scale and how to fix it ยท You have improved real cluster throughput via communication redesign ยท You can trace a distributed hang across ranks and identify the root cause ยท You are comfortable working at the boundary between hardware and runtime, Masterโ€™s, or Bachelorโ€™s + 1 year of relevant experience.

Benefits & conditions

This position is eligible for visa sponsorship. Benefits Include *Comprehensive medical, dental, and vision benefits *Bonus *401K Plan *Generous paid time off, sick leave and holidays *Paid Parental Leave *Employee Assistance Program *Life insurance and disability FULL_TIME Organization https://startup.jobs/logos/42630 Institute of Foundation Models None Startups Place

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role โ€” technically off-topic, practically not.

2:51 min

Designing hardware infrastructure and networking for distributed compute

Anshul Jindal Anshul Jindal +1 ยท World Congress 2025

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues ยท World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balรกzs Kiss ยท World Congress 2023

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters ยท World Congress 2023

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu ยท Europe 2026 Virtual

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 ยท World Congress 2026 Europe

Videos

See all

Related articles

See all