> Markdown version of [/jobs/ext/2271047-technical-program-manager-local-ai-agents](https://www.wearedevelopers.com/jobs/ext/2271047-technical-program-manager-local-ai-agents). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Technical Program Manager -Local AI Agents - **Company:** NVIDIA Ltd. - **Location:** Redmond, WA, United States - **Experience:** Expert - **Salary:** $168,000.0 - $258,750.0 - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Artificial Intelligence, Application Integration Architecture, C++ (Programming Language), Program Optimization, Nvidia CUDA, Computer Engineering, Computer Literacy, Desktop Computing, Programming Tools, Device Drivers, Design of User Interfaces, Human-Computer Interaction, Python (Programming Language), Microsoft Software, Open Source Technology, Release Management, Privacy Controls, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Multi-Agent Systems, Software Security, Information Technology, Low Latency, ONNX (Open Neural Network Exchange) Format, TensorRT, Virtual Agents, Multiplatform, Programming Languages - **Published:** August 27, 2026 - **Apply:** https://www.careerbuilder.com/job-details/technical-program-manager-local-ai-agents-redmond-wa--ae718b39-5403-47ae-aa3b-3c962abe2a3a ## About the Role * Bachelor's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience in practice. * 8+ years leading complex software, platform, systems, or AI/ML programs from concept through production release. * Technical proficiency in contemporary AI software stacks, encompassing model prediction, agent coordination, tool application, evaluation, and local or hybrid deployment. * Experience building coordinated plans across multiple engineering organizations and resolving technical dependencies without direct authority. * Experience defining measurable quality and release criteria, using data to make tradeoffs, and separating prototypes from validated capabilities. * Ability to explain architecture, risk, and program status clearly to technical teams, partners, and executives. Ways to stand out from the crowd: * Hands-on experience with Python, C++, or software automation. * Familiarity with OpenClaw, Hermes, LangChain or similar agent frameworks; MCP; and persistent memory, skills, multi-agent, or computer-use patterns. * Familiarity with local inference technologies such as TensorRT-LLM, Ollama, llama.cpp, vLLM, PyTorch, ONNX Runtime, or Windows ML, plus model optimization or quantization. * Experience with Windows systems, CUDA or GPU acceleration, sandboxing, policy controls, privacy-aware networking, or secure credential handling. * Experience collaborating with Microsoft, OEMs, ISVs, open-source maintainers, or domain-solution partners., Application Integration, Artificial Intelligence (AI), Artificial Intelligence (AI) Agents, Artificial Intelligence (AI) Programming Languages, CUDA (Compute Unified Device Architecture), Communication Skills, Computer Engineering, Computer Science, Cross-Functional, Customer/Client Research, Desktop PC, Device Drivers, Documentation, Ecosystems, GPU (Graphics Processing Unit), MCP - Microsoft Certified Professional, Memory Hardware, Microsoft Product Family, Microsoft Windows Operating System, Multiplatform/Cross-Platform, OEM (Original Equipment Manufacturer), Onboarding, Open Source, Predictive Modeling, Privacy Controls, Programming Tools, Project/Program Coordination, Project/Program Management, Prototyping, Public/Media/Press/Analyst Relations, Quality Metrics, RTX, Reporting Dashboards, Risk, Technical Leadership, User Interface/Experience (UI/UX) ## Description We seek a Senior Technical Program Manager to lead multi-functional initiatives involving local and hybrid AI agents. You will unite engineering, research, product, security, developer relations, and external collaborators behind a clear roadmap and measurable results. Your responsibilities include local model inference, agent runtimes, Windows integration, developer tools, security controls, and user experiences. Additionally, you will gather insights from open ecosystems like Hermes and OpenClaw, desktop-agent projects such as Perplexity, and domain-specific agents to guide NVIDIA's focus on developer and user priorities. What you'll be doing: * Own the coordinated roadmap from prototype through release for local AI agent capabilities, with clear scope, achievements, owners, dependencies, and completion criteria. * Translate product goals and ecosystem signals into harmonized plans covering agent frameworks, MCP and tool connections, memory and skills, computer use, local inference, model routing, and application integration. * Partner with engineering and research teams to define evaluation and release gates for task success, latency, efficiency, memory use, power, reliability, setup time, privacy, and user trust. * Coordinate programs across Windows platform groups, GPU and driver groups, model and runtime groups, product security, developer relations, Microsoft, OEMs, ISVs, and open-source communities. * Build concise dashboards and decision forums that surface program health, technical tradeoffs, risks, and evidence; communicate clearly with engineers and senior leaders. * Use developer and customer feedback to improve onboarding, documentation, samples, compatibility, and adoption across supported NVIDIA systems. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [On a Secret Mission: Developing AI Agents](https://www.wearedevelopers.com/videos/1510-on-a-secret-mission-developing-ai-agents) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)