> Markdown version of [/jobs/ext/2695744-principal-engineer-in-dgx-cloud](https://www.wearedevelopers.com/jobs/ext/2695744-principal-engineer-in-dgx-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Engineer in DGX Cloud - **Company:** NVIDIA Corporation - **Location:** Santa Clara, CA, United States - **Salary:** $272,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Nvidia CUDA, Computer Programming, Distributed Systems, Python (Programming Language), Prometheus, Software Engineering, System Programming, Systems Integration, Workflow Management Systems, Scripting, Graphics Processing Unit (GPU), Google Cloud, Grafana, Containerization, Kubernetes, Infrastructure Automation Frameworks, Data Management, Slurm, Data Pipelines, Docker, Programming Languages - **Published:** September 3, 2026 - **Apply:** https://www.careerbuilder.com/job-details/principal-software-engineer-dgx-cloud-santa-clara-ca--eb9ddd4a-69a0-4c31-a014-86689cbaf5f4 ## About the Role * 16+ years of progressive industry experience * Master's or Bachelor's degree, or equivalent experience defining and shipping complex distributed systems. * Deep, hands-on expertise in establishing, operating, and scaling services in a fast paced, high-reliability environment. * Thrive in ambiguous, fast paced environments by rapidly testing ideas, iterating toward working solutions, and then hardening the winners into reliable, scalable systems. * Outstanding proficiency in modern systems programming languages such as Go, Java, or Python. * Proven track record of defining, owning, and evolving the architecture of high-scale distributed systems, including advanced patterns for APIs, control planes, and data pipelines. * Deep understanding of global cloud infrastructure (AWS, GCP, Azure) and container ecosystems (Docker, Kubernetes). * Demonstrated ability to drive technical strategy and influence outcomes across organizational boundaries. * Outstanding ability to communicate complex technical concepts, drive organizational consensus, and mentor high-performing engineers. Ways to Stand Out from the Crowd: * A history of successfully leading the development and adoption of organization-wide workflow orchestration systems for petabyte-scale infrastructure. * Experience in a Principal/Staff+ capacity, delivering measurable improvements in operational efficiency, reliability, and security across a large engineering org. * Deep familiarity with the operational and deployment aspects of the NVIDIA AI/ML software stack (CUDA, cuDNN, containerization). * Patent contributions or a strong publication record in areas related to distributed systems, cloud computing, or infrastructure automation., Adoption, Amazon Web Services (AWS), Application Programming Interface (API), Artificial Intelligence (AI), Automation, Business Processes, Business Solutions, CUDA (Compute Unified Device Architecture), Cloud Computing, Coaching, Communication Skills, Customer Experience, Customer Relations, Data Management, Distributed Computing, Docker, Ecosystems, GCP (Good Clinical Practices), GPU (Graphics Processing Unit), High Reliability, Java, Leadership, Mentoring, Microsoft Windows Azure, Operational Improvement, Operational Measurement, Operational Strategy, Patents, Programming Languages, Publications, Python Programming/Scripting Language, Software Engineering, System Integration (SI), Systems Scalability, Systems/Internals Programming, Technical Leadership, Technical Strategy, Vehicle Fleets ## Description * Lead the build and development of next-generation APIs, state management, and workflow orchestration systems that automate fleet lifecycle operations at a massive scale. * Drive technical alignment across dependent systems and partner teams to ensure cohesive integration, clear interfaces, and reliable end-to-end workflows, with a strong focus on delivery. * Act as a force-multiplier by coaching, mentoring, and encouraging senior engineers, elevating the technical standards and guidelines across the organization. * Maintain an incredible focus on the customer experience and product requirements, translating deep technical insight into high-impact business solutions. * Partner with executive and engineering leadership to codify critical business processes into self-measuring, scalable, and operationally consistent platforms, drastically reducing manual toil. * Direct the integration strategy for key technologies, including common AI schedulers (e.g., Kubernetes, Slurm) and innovative observability systems (e.g., Prometheus, OpenTelemetry, Grafana). ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai)