> Markdown version of [/jobs/ext/1234120-senior-devops-engineer-platform-engineering](https://www.wearedevelopers.com/jobs/ext/1234120-senior-devops-engineer-platform-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior DevOps Engineer, Platform Engineering - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $176,000.0 - $276,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Business Analytics Applications, Microsoft Azure, Computer Engineering, Data Centers, DevOps, Distributed Systems, Github, Python (Programming Language), Linux System Administration, Ansible, Prometheus, Azure Machine Learning, Management of Software Versions, Graphics Processing Unit (GPU), Google Cloud, Grafana, HybridCloud, Gitlab, Git Flow, Kubernetes, Information Technology, Hardware Infrastructure, Terraform, Network Server, Jenkins - **Published:** July 11, 2026 - **Apply:** https://www.disabledperson.com/jobs/73606188-senior-devops-engineer-platform-engineering ## About the Role * BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience, with over 6+ years of relevant industry background. * Advanced skills in Python for scripting, tooling, and automation. * Deep expertise with Kubernetes, Helm, and container orchestration in production environments. * Verified background in building and maintaining CI/CD pipelines at scale (Jenkins, GitHub/GitLab Actions and Runners, or similar). * Solid understanding of Linux systems administration, networking, and distributed systems. * Experience with release engineering practices including semantic versioning, release gating, and change management. * Hands-on experience with observability stacks (Prometheus, Grafana, ELK, or similar). Ways to stand out from the crowd: * Experience with GPU infrastructure and AI/ML platform engineering at scale. * Background in BareMetal and hybrid cloud (AWS, GCP, Azure) environment management. * Familiarity with NVIDIA Metropolis, DeepStream, or similar AI video analytics platforms. * Experience with GitOps workflows, Infrastructure as Code (Terraform, Ansible). * Track record of driving DevOps culture transformation and developer experience improvements. ## Description * Compose, build, and maintain scalable CI/CD pipelines using Jenkins, GitHub/GitLab Actions and Runners for Metropolis software products. * Develop and manage Kubernetes-based platform infrastructure supporting AI/ML workloads on NVIDIA Data Center GPUs. * Build and implement scaling and performance measurement frameworks within Kubernetes to ensure platform reliability and efficiency under AI/ML workload demands. * Define and implement release engineering processes, branching strategies, versioning standards, and gating criteria. * Drive developer efficiency by building and maintaining DevOps MCP servers, tooling, and automation frameworks. * Own observability and monitoring infrastructure using Prometheus, Grafana, and log aggregation pipelines. * Troubleshoot hardware and operating system issues across BareMetal and GPU-accelerated servers to minimize downtime and maintain platform stability. ## Related Videos - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)