> Markdown version of [/jobs/ext/1919308-senior-software-engineer-distributed-systems](https://www.wearedevelopers.com/jobs/ext/1919308-senior-software-engineer-distributed-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer - Distributed Systems - **Company:** NEUROSPARK, LLC - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $170,000.0 - $350,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Computer Clusters, Software Debugging, Distributed Systems, Graphics Processing Unit (GPU), Load Balancing, Large Language Models, Concurrency, Build Management, Kubernetes, Low Latency, Free and Open-Source Software, Hardware Infrastructure - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=759dfa70e92e6472 ## About the Role This is a senior individual-contributor role. We're looking for someone who has built systems like this before and can operate independently from the first week. * Substantial experience building and operating large-scale distributed systems in production - you've owned something load-bearing, not just contributed to it * Track record of designing core systems from zero to one, and living with the consequences of your architectural decisions * Hands-on experience with scheduling, load balancing, request routing, or resource allocation systems * Strong systems fundamentals - operating systems, networking, concurrency - and the ability to reason quantitatively about system behavior using queueing theory, control theory, or similar * Fluency in a performance-sensitive language (Go, Rust, or C++), with the profiling and optimization instincts that come from actually chasing latency in production * Comfort with GPU infrastructure and LLM inference fundamentals - batching, KV cache behavior, throughput/latency tradeoffs; deep expertise here is a plus, but strong distributed-systems judgment matters more * Clear technical writing - you can make a hard design decision legible to people who weren't in your head * An AI-native way of working - you use AI tools daily and have your own view of how they change how infrastructure gets built Preferred Qualifications * Mandarin proficiency is a plus Nice to have: Kubernetes and multi-cloud operations experience; open-source contributions to inference, serving, or scheduling projects; experience operating GPU clusters at scale. ## Description Serving inference at scale is a scheduling problem. Requests arrive with wildly different shapes and latency expectations, GPUs are heterogeneous and expensive, and the difference between a platform that's fast and one that's economical usually comes down to how well work gets placed. That system is what you'll own. You'll design and build the scheduling and routing layer of our platform: how requests get admitted, prioritized, batched, and placed across a heterogeneous multi-cloud GPU fleet, under real multi-tenant load and real latency commitments. This is core-systems work with a clean slate - you'll be making the foundational architectural decisions, not maintaining someone else's, and the quality of those decisions will show up directly in our margins and our customers' latency numbers. You'll work close to the metal and close to the math. Some days that means reasoning about queueing behavior and control loops on a whiteboard; other days it means profiling Go or Rust until the tail latency comes down. We're a small team, so you'll own systems end-to-end - design, implementation, rollout, and the production reality afterward. Responsibilities * Own the scheduling and routing layer - design and build request admission, prioritization, batching, and placement across a heterogeneous GPU fleet spanning multiple clouds and accelerator types * Engineer for latency and utilization at once - drive down tail latency while driving up fleet utilization; these fight each other, and resolving that tension well is the job * Model the system, not just code it - apply queueing theory, control theory, and load-shedding principles to make the platform behave predictably under bursty, multi-tenant traffic * Build multi-tenant fairness and isolation - ensure priority guarantees and SLO commitments hold when the fleet is saturated and customers are competing for the same capacity * Own it in production - instrument, observe, and debug distributed behavior in a live system; carry your designs through rollout and real-world load * Set the technical bar - make foundational architecture decisions, write the design docs that anchor them, and raise the engineering standard of everyone around you ## Related Videos - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Concurrency with Go](https://www.wearedevelopers.com/videos/191-concurrency-with-go) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)