> Markdown version of [/jobs/ext/2846144-software-development-engineer-ai-ml-networking-disaggregated-inference-annapurna-labs-elastic-c](https://www.wearedevelopers.com/jobs/ext/2846144-software-development-engineer-ai-ml-networking-disaggregated-inference-annapurna-labs-elastic-c). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Development Engineer - AI/ML Networking Disaggregated Inference, Annapurna Labs, Elastic C - **Company:** Annapurna Labs - **Location:** Cupertino, CA, United States - **Experience:** Experienced - **Salary:** $165,200.0 - $223,600.0 - **Contract:** Internship / Graduate position - **Skills:** Artificial Intelligence, Amazon Web Services, Systems Engineering, Code Review, Computer Programming, Software Design Patterns, Linux, InfiniBand, Remote Direct Memory Access, Software Engineering, Large Language Models, Information Technology, Build Process, Machine Learning Operations, Software Coding, Network Server, Software Version Control - **Published:** September 11, 2026 - **Apply:** https://www.careerboard.com/us/en/find-jobs-in-United-States/-4AA953068A3843D30F/ ## About the Role Strong C/C+ and a genuine interest in low-level, performance-critical systems - solid command of Linux, memory, and writing fast code. The instinct to ask "how fast could this go?" and the discipline to measure it. Exposure to high-speed networking, HPC interconnects, or GPU/accelerator systems (RDMA, InfiniBand, libfabric, UCX, NCCL, MPI) is a strong plus; embedded-systems experience is welcome. Prior AI/ML experience is not required - if you're a strong systems engineer eager to learn, we'll teach you the ML side. If you like solving genuinely hard problems, working alongside HPC and ML customers, iterating fast, and shipping at a scale few places can offer, come join us. You'll work alongside senior engineers and Principal Engineers who've built this layer from the ground up, with real room to grow your scope and technical depth - on a team at the leading edge of AI/ML infrastructure., BASIC QUALIFICATIONS - 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Experience programming with at least one software programming language - Experience with C/C+ PREFERRED QUALIFICATIONS - 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Bachelor's degree in computer science or equivalent ## Description We're looking for an engineer to work at the frontier of disaggregated inference: splitting LLM serving into separate prefill and decode pools and moving the model's KV cache between them at the limit of what the hardware allows. Get it right and users get answers in milliseconds; get it wrong and the fastest accelerators in the world sit idle waiting on data. You'll build components of the high-speed transfer path that make that difference, and you'll learn to measure success in how close we run to the theoretical peak of the machine. In this role you will: Build and optimize the low-level data-movement software that transfers KV cache and activations across accelerators, servers, and heterogeneous memory - over AWS's highest-performance network fabric. Profile real workloads, find the true bottleneck, and close the gap between "it works" and "it runs fast" - pushing components toward the hardware's limit. Work across the stack - from network transport up to the inference frameworks - learning from the teams building the chips, runtime, and models. Deliver features that ship to our largest clusters, for our largest customers, serving the largest AI models in production. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Build a CI/CD pipeline to automate code reviews and ensure code quality](https://www.wearedevelopers.com/videos/349-build-a-ci-cd-pipeline-to-automate-code-reviews-and-ensure-code-quality) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters](https://www.wearedevelopers.com/magazine/571-dev-digest-162-ai-careers-mcp-aws-best-practices-floppy-sweaters) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)