> Markdown version of [/jobs/ext/1930587-senior-software-engineer-infrastructure-automation-and-distributed-systems](https://www.wearedevelopers.com/jobs/ext/1930587-senior-software-engineer-infrastructure-automation-and-distributed-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Infrastructure Automation and Distributed Systems - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $224,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Linux, Distributed Systems, Perl (Programming Language), Python (Programming Language), OpenStack, Ruby, Cloud Platform System, Multi-Cloud, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Bare Metal, Slurm, Docker, Golang - **Published:** August 5, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/occis9gpwk ## About the Role * BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics) or equivalent experience. * 12+ years of relevant experience. * A track record showing a good balance between initiating your own projects, convincing others to collaborate with you and collaborating well on projects initiated by others. * Experience with infrastructure automation and distributed systems design developing tools for running large scale private or public cloud system in production. * Experience in one or more of the following: Python, Go, Perl or Ruby. * In depth knowledge in one or more of Linux, Networking, Storage, and Containers. Ways To Stand Out From The Crowd * Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive. Experience accelerating positive impact to the business using coding assistant(s), MCP servers, or AI agents. * Experience working with or developing bare metal as a service (BMaaS) associated systems. * Experience working with or developing multi-cloud infrastructure services and running private or public cloud systems based on one or more of Kubernetes, OpenStack, Docker or Slurm. * Experience teaching reliability (e.g SRE) or more general cloud systems good practices to peers or to other companies (e.g CRE). * Background with NVIDIA Collective Communication Library (NCCL). No prior experience having worked in a team of any particular name or having worked in a ML/AI focused team are required but also a nice to have. ## Description * Design, build, deploy, and run infrastructure services & manage the software life cycle in scope to meet our business goals. * Participate in the definition of our internal facing service level objectives and error budgets as part of our overall observability strategy. * Eliminate toil or automate it where the ROI of building and maintaining automation is worth it. * Practice sustainable blameless incident prevention and incident response while being a member of an oncall rotation. * Consult with and provide consultation for peer teams on systems design best practices. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)