> Markdown version of [/jobs/ext/2298155-senior-manager-software-engineering-agentic-it-operations](https://www.wearedevelopers.com/jobs/ext/2298155-senior-manager-software-engineering-agentic-it-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Manager, Software Engineering - Agentic IT Operations - **Company:** NVIDIA Corporation - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $248,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Code Review, Continuous Delivery, Continuous Integration, DevOps, Monitoring of Systems, Information Technology Operations, Python (Programming Language), Automation of Marketing, Reliability Engineering, Prometheus, Software Engineering, Software Systems, Systems Integration, Datadog, Large Language Models, Grafana, Containerization, Kubernetes, Production Code, Build Process, Data Pipelines, Pagerduty, Servicenow - **Published:** August 29, 2026 - **Apply:** https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Manager--Software-Engineering---Agentic-IT-Operations_JR2023714 ## About the Role We are seeking a hands-on technical leader to build and lead a high-performance engineering organization that architects, delivers, and operates production-grade software systems at global scale. You will be responsible for transforming enterprise IT operations from manual, reactive workflows into fully automated, AI-driven platforms that scale with NVIDIA's hyper-growth. This role demands deep software engineering expertise, systems thinking, and the ability to drive large-scale technical transformation with measurable business outcomes. Exceptional interpersonal, written, and verbal communication skills are vital for success., * Bachelor's or Master's degree in a related field, or equivalent experience * 10+ overall years of hands-on software engineering experience, with deep expertise in at least one of: Infrastructure, SRE, DevOps, or Production Engineering. 5+ years leading engineering teams, with direct experience hiring, growing, and managing IT engineers. * Demonstrated ability to build engineering teams from zero and scale them in a high-growth, high-ambiguity environment. * Deep expertise in designing and shipping production software systems-including integrations, automation platforms, and data pipelines-for complex enterprise operations at scale. * Track record of modernizing enterprise IT operations platforms (e.g., asset management, endpoint services, IT supply chain, infrastructure operations) and deploying agentic AI into production-including multi-step autonomous execution, human-in-the-loop safeguards, exception handling, and governance frameworks with measurable business outcomes. * Production-grade proficiency with infrastructure-as-code, CI/CD, containerization (Kubernetes), and cloud platforms (AWS, GCP, or Azure). * Experience with monitoring and observability tools (Prometheus, Grafana, Datadog, PagerDuty, or similar). * Fluent in Python, Go, or equivalent languages-able to architect, write, and review production-quality code, not just scripts. * Executive-level communication skills with the ability to influence technical direction across engineering, product, and senior leadership. * Proven ability to translate complex technical capabilities into quantifiable business value and present to VP/C-level audiences. ## Description As a Senior Engineering Manager, you will own the technical vision, architecture, and delivery of enterprise-scale automation platforms that eliminate manual workflows and enable NVIDIA to operate at 10x scale without proportional headcount growth. You will build and lead a team of IT engineers, set the technical bar, and personally contribute to system design and code. Core responsibilities include: * Architect and ship agentic AI systems using LLM-based agents, tool calling, RAG, and orchestration frameworks delivering production-grade AI-assisted operations across enterprise IT domains including employee support, endpoint services, and IT support operations. * Design and deploy autonomous AI agents that execute complex, multi-step enterprise workflows end-to-end coordinating approvals, vendor handoffs, cross-system data reconciliation, and exception handling with human-in-the-loop controls delivering measurable improvements in availability, cycle time, cost, and compliance. * Engineer robust integration and automation platforms spanning ServiceNow, ERP and procurement systems, endpoint-management platforms, Own the full stack infrastructure, data pipelines, APIs, and user-facing applications. * Set the engineering standard through hands-on technical leadership co-authoring production code, conducting rigorous code reviews, and personally driving system design for the most critical components. * Recruit, develop, and retain top-tier engineering talent. Build a high-performing team culture grounded in engineering excellence, ownership, and continuous delivery. * Define and execute a multi-quarter technical roadmap for automation and agentic operations across enterprise IT, with each initiative tied to quantifiable business outcomes (cost reduction, throughput, SLA improvement, headcount avoidance). * Drive disciplined execution-project prioritization, milestone tracking, capacity planning, and on-time delivery-while maintaining engineering velocity in a fast-moving environment. * Own talent strategy for the team, including hiring pipelines, performance calibration, and career development that builds a deep bench of engineering leaders. ## Related Videos - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)