> Markdown version of [/jobs/ext/1488676-senior-system-software-engineer-agentic-inference-dynamo](https://www.wearedevelopers.com/jobs/ext/1488676-senior-system-software-engineer-agentic-inference-dynamo). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior System Software Engineer, Agentic Inference - Dynamo - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $224,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Computer Engineering, Software Debugging, Python (Programming Language), Open Source Technology, Software Engineering, System Software, Graphics Processing Unit (GPU), Large Language Models, Caching, Information Technology, Free and Open-Source Software, Front End Software Development, TensorRT - **Published:** July 29, 2026 - **Apply:** https://www.disabledperson.com/jobs/73912907-senior-system-software-engineer-agentic-inference-dynamo ## About the Role * Masters or PhD or equivalent experience * 10+ years in Computer Science, Computer Engineering, or related field * Ability to work in a fast-paced, agile team environment * Excellent Rust/Python programming and software design skills, including debugging, performance analysis, and test design. * Understanding of modern LLM API semantics, including structured outputs, tool calling, reasoning controls, token accounting, context management, and multimodal inputs. Ways to stand out from the crowd: * Prior contributions to open-source AI inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang). * Experience optimizing GPU memory, KV and prefix caches, or high-performance networking for long-context, reasoning, and tool-calling workloads. * Understanding of LLM-specific inference challenges for agentic workloads, including context and reasoning-token growth, bursty tool-call-driven traffic, multi-turn state reuse, and scheduling across concurrent trajectories. * Prior experience integrating self-hosted LLM serving stacks with agent harnesses such as OpenCode, Codex, Claude Code, and Pi, including compatibility for APIs, streaming, structured outputs, tool calls, and session semantics. ## Description * In this role, you will develop open source software to serve inference of trained AI models running on GPUs. * Contribute to the development of disaggregated serving for Dynamo-supported inference engines (vLLM, SGLang, TRT-LLM) and expand these capabilities to support agentic inference workloads, including long-horizon reasoning, tool calling, and stateful, multi-turn execution. * Innovate in inference-state management for long-running agents, including KV- and prefix-cache reuse and transfer across heterogeneous memory and storage hierarchies with NIXL, to reduce repeated prompt processing, improve latency and token throughput, maximize GPU utilization, and lower per-token and per-task costs for self-hosted LLMs. * Build and evolve Dynamo's distributed inference frontend across vLLM, SGLang, and TensorRT-LLM, delivering day-0 support for new models, model-specific request parameters, upstream API compatibility, and stateful Responses API semantics. * Balance a variety of objectives: build robust, scalable, high performance software components to support our distributed inference workloads; work with team leads to prioritize features and capabilities; load-balance asynchronous requests across available resources; optimize throughput under latency constraints; and integrate the latest open source technology. ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Flexibility is Key: Unlocking the Advantages of Versatile Software Solutions for Strategic Innovatio](https://www.wearedevelopers.com/videos/1943-flexibility-is-key-unlocking-the-advantages-of-versatile-software-solutions-for-strategic-innovatio) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Architecture Communication Canvas](https://www.wearedevelopers.com/videos/674-architecture-communication-canvas) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)