> Markdown version of [/jobs/ext/1133406-senior-technical-program-manager](https://www.wearedevelopers.com/jobs/ext/1133406-senior-technical-program-manager). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Technical Program Manager - **Company:** NVIDIA Ltd. - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $168,000.0 - $322,000.0 - **Contract:** Permanent contract - **Skills:** Big Data, Cloud Computing, Distributed Systems, SQL Databases, AI Infrastructure, Business Intelligence Development Studio, Grafana, Apache Spark, Mttr, Kubernetes, Data Analytics, Databricks - **Published:** July 2, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=74bd72fa9eed925e ## About the Role * 8+ years of Technical Program Management experience, with at least 3 years in infrastructure, platform, or reliability-focused domains. * Strong hands-on data analytics skills - comfortable writing SQL, working with large telemetry datasets, and building dashboards (Grafana, Superset, Databricks, or equivalent). * Demonstrated ability to define and operationalize reliability metrics (MTBI, MTTR, availability SLAs) and drive engineering teams toward measurable improvements. * Proven ability to lead deep-dive investigations across ambiguous, multi-system problems and translate findings into long-term solutions. * Excellent executive communication skills - able to distill complex technical findings into clear, decision-ready narratives for senior leadership. * MS in EE, CS, or equivalent experience. Ways to stand out from the crowd: * Familiarity with NVIDIA GPU architectures and DGX/HGX infrastructure. * Experience with Databricks, Apache Spark, or other large-scale data processing platforms. * Hands-on experience with Grafana, Superset, or similar observability/BI tooling. * Background in cloud-native infrastructure, Kubernetes, or large-scale distributed systems. ## Description As a Senior Technical Program Manager with a passion for data-driven operations, you will lead the DGX Cloud Fleet Health reporting program - delivering real-time, actionable insights on the availability and reliability of our GPU fleet. A core focus of this role is advancing Mean-Time-Between-Interruption (MTBI): understanding the root causes of fleet interruptions, surfacing patterns in the data, and driving cross-functional programs to measurably extend fleet uptime. You will partner closely with Capacity Operations, Infrastructure, SRE, and Engineering teams to translate complex fleet signals into decisions that directly improve customer experience. Join us in making a significant impact on the world's most powerful AI infrastructure. What You'll Be Doing: * Define and own the metrics framework for measuring fleet health, reliability, and MTBI across a diverse and rapidly scaling GPU fleet. * Lead hands-on data investigations - querying telemetry, correlating failure signals, and building statistical models - to identify the root causes of interruptions and quantify their impact. * Own and drive execution of cross-functional MTBI improvement programs end-to-end - from translating analytical findings into a prioritized roadmap, to holding teams accountable to milestones and delivering measurable reliability gains. * Build and maintain dashboards, automated anomaly detection, and alerting frameworks that surface gaps in fleet health reporting in real time. * Anticipate and close reporting gaps with new cloud providers and hardware platforms by working closely with Infrastructure bring-up teams. * Communicate complex data findings and program status clearly to senior leadership, turning raw signals into crisp narratives and recommendations. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Should Tech Managers Be Developers First? Pros and Cons](https://www.wearedevelopers.com/magazine/327-should-tech-managers-be-developers-first-pros-and-cons) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)