> Markdown version of [/jobs/ext/1953286-software-engineer-cuda-deep-learning-systems](https://www.wearedevelopers.com/jobs/ext/1953286-software-engineer-cuda-deep-learning-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, CUDA Deep Learning Systems - **Company:** NVIDIA Ltd. - **Location:** Austin, TX, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Distributed Systems, Python (Programming Language), Node.Js, Open Source Technology, Performance Tuning, System Programming, Graphics Processing Unit (GPU), Deep Learning - **Published:** August 6, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/software-engineer-cuda-deep-learning-systems-austin-tx-usa-58815184 ## About the Role _ to identify bottlenecks in training and inference * Collaborate with AI researchers, HW/SW architects, kernel and compiler experts, and CUDA drivers * Develop exploratory tools and runtime systems to profile and accelerate new DL paradigms * Write clean, maintainable code to enable prototypes to transition to open-source releases or products Tasks * BS/MS/PhD in CS, CE, EE, or related field (or equivalent experience) * 2+ years of relevant industry or academic experience * Strong proficiency in C++ and Python * Solid fundamentals in Deep Learning with a focus on transformers * Strong understanding of distributed computing, multi-node scaling, and performance challenges in cluster environments * Proven experience in systems programming, computer architecture, and low-level performance optimization * Hands-on experience with CUDA programming, kernel optimization, and workload profiling Key requirements * equity * benefits * remote/hybrid options * competitive compensation ## Description Experteer Overview In this role you will advance deep learning workloads by optimizing CUDA-based systems for cutting-edge AI models. You will work with a cross-functional team to prototype high-performance kernels and distributed pipelines that scale from a single node to clusters. The job blends research and practical implementation, aiming to maximize accelerator utilization and memory bandwidth across training and inference. This is a chance to shape next-generation AI systems on modern GPUs and contribute to open-source and internal tooling. You will join a highly technical, research-oriented group tackling uncharted optimization and architecture challenges. Compensation / Benefits * Explore and prototype system optimizations at the intersection of high-level DL frameworks and CUDA * Architect and optimize distributed computing systems from single-node to cluster-scale * Design, implement, and optimize custom high-performance CUDA kernels * Analyze hardware-software interactions to identify bottlenecks in training and inference * Collaborate with AI researchers, HW/SW architects, kernel and compiler experts, and CUDA drivers * Develop exploratory tools and runtime systems to profile and accelerate new DL paradigms * Write clean, maintainable code to enable prototypes to transition to open-source releases or products Tasks * BS/MS/PhD in CS, CE, EE, or related field (or equivalent experience) * 2+ years of relevant industry or academic experience * Strong proficiency in C++ and Python * Solid fundamentals in Deep Learning with a focus on transformers * Strong understanding of distributed computing, multi-node scaling, and performance challenges in cluster environments * Proven experience in systems programming, computer architecture, and low-level performance optimization * Hands-on experience with CUDA programming, kernel optimization, and workload profiling Key requirements * equity * benefits * remote/hybrid options * competitive compensation ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [CUDA Python: GPU programming for the modern developer](https://www.wearedevelopers.com/videos/100221-cuda-python-gpu-programming-for-the-modern-developer) - [The weekly developer show: Boosting Python with CUDA, CSS Updates & Navigating New Tech Stacks](https://www.wearedevelopers.com/videos/1293-the-weekly-developer-show-boosting-python-with-cuda-css-updates-navigating-new-tech-stacks) ## Related Articles - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs](https://www.wearedevelopers.com/magazine/557-dev-digest-157-cuda-in-python-gemini-code-assist-and-back-dooring-llms) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)