> Markdown version of [/jobs/ext/1437360-ai-and-ml-infra-software-engineer-gpu-clusters-new-college-grad-2026](https://www.wearedevelopers.com/jobs/ext/1437360-ai-and-ml-infra-software-engineer-gpu-clusters-new-college-grad-2026). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI and ML Infra Software Engineer, GPU Clusters - New College Grad 2026 - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Starter - **Salary:** $124,000.0 - $195,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, Computer Clusters, DevOps, General Parallel File Systems, InfiniBand, Python (Programming Language), Machine Learning, Data Processing, Google Cloud, Cloud Platform System, High Performance Computing, Pytorch, System Availability, Delivery Pipeline, Parallel Computation, Kubernetes, Information Technology, Slurm, Machine Learning Operations, Docker, Golang - **Published:** July 25, 2026 - **Apply:** https://www.disabledperson.com/jobs/73846004-ai-and-ml-infra-software-engineer-gpu-clusters-new-college-grad-2026 ## About the Role * Recent graduate with a MS, PhD or equivalent experience in Computer Science or related field, with proven experience in AI/ML and HPC workloads and infrastructure. * Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure as well as in-depth knowledge of accelerated computing (e.g., GPU, custom silicon), storage (e.g., Lustre, GPFS, BeeGFS), scheduling & orchestration (e.g., Slurm, Kubernetes, LSF), high-speed networking (e.g., Infiniband, RoCE, Amazon EFA), and containers technologies (Docker, Enroot). * Expertise in running and optimizing large-scale distributed training workloads using PyTorch (DDP, FSDP), NeMo, or JAX. Also, possess a deep understanding of AI/ML workflows, encompassing data processing, model training, and inference pipelines. * Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure) in addition to experience with parallel computing frameworks and paradigms. * Passion for continual learning and keeping abreast of new technologies and effective approaches in the AI/ML infrastructure field. * Excellent communication and collaboration skills, with the ability to work effectively with teams and individuals of different backgrounds. ## Description * Collaborate closely with our AI and ML research teams to understand their infrastructure needs and obstacles, translating those observations into actionable improvements. * Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization. * Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results. * Collaborate with diverse teams, including researchers, data engineers, and DevOps professionals, to build a seamless and coordinated AI/ML infrastructure ecosystem. * Stay on top of the latest advancements in AI/ML technologies, frameworks, and effective strategies, and promote their implementation within the company. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)