> Markdown version of [/jobs/ext/1387840-software-engineer-ai-and-dl-kernel-libraries](https://www.wearedevelopers.com/jobs/ext/1387840-software-engineer-ai-and-dl-kernel-libraries). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, AI and DL Kernel Libraries - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Computer-Aided Design, Application Programming Interfaces (APIs), Artificial Intelligence, Apache HTTP Server, Computer Vision, C++ (Programming Language), Compilers, Code Generation, Program Optimization, Profiling, Nvidia CUDA, Computer Programming, Python (Programming Language), Linux Kernel, Machine Learning, Open Source Technology, Performance Tuning, Recommender Systems, Tensorflow, Software Engineering, Systems Architecture, System Programming, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Deep Learning, Generative AI, Gpu Programming, Information Technology, ONNX (Open Neural Network Exchange) Format, TensorRT - **Published:** July 22, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/p6l9up47k3 ## About the Role * Master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience. * 3+ years of relevant industry, research, or systems software development experience in machine learning, deep learning systems, compilers, or GPU software. More experience is expected for senior-level candidates. * Strong programming skills in C/C++ and Python, with hands-on experience developing high-performance software. * Solid experience with CUDA development and GPU programming fundamentals. * Strong experience developing or using deep learning frameworks such as PyTorch, JAX, TensorFlow, or ONNX. * Good understanding of linear algebra, performance analysis, profiling, and code optimization. * Experience designing software abstractions, APIs, or higher-level system architecture for performance-sensitive systems. * Familiarity with modern machine learning and inference system trends, especially around LLMs and generative AI. * For senior candidates, strong experience in GPU kernel development and performance optimization, especially using CUDA C/C++, cuTile, Triton, or similar technologies, is expected. Ways To Stand Out From The Crowd * Hands-on experience with inference engines and runtimes such as vLLM, SGLang, MLC, TensorRT-LLM, or similar systems. * Background in domain-specific compiler, code generation, or library solutions for LLM inference and training. * Expertise in machine learning compilers or IR systems such as MLIR, Apache TVM, TensorIR, or related technologies. * Practical experience with GPU performance modeling, computer architecture, or accelerator-oriented software design. * Open-source project ownership or meaningful contributions in deep learning systems, compilers, kernels, or inference infrastructure. , , JR2019913 ## Description * Develop production-quality software that ships as part of NVIDIA's AI software stack, including cuDNN, FlashInfer, and optimized support for large language model inference workloads. * Innovate and develop new AI systems technologies for efficient inference, with a focus on performance, scalability, maintainability, and usability. * Design, implement, and optimize kernels for high-impact AI workloads across LLM inference, generative AI, computer vision, autonomous driving, and recommender systems. * Design and implement extensible software abstractions for deep learning libraries, LLM serving engines, and runtime systems. * Build and improve just-in-time compilation, code generation, and runtime technologies for performance-critical GPU workloads. * Analyze workload performance, tune current software, and propose improvements to future software and hardware-software interfaces. * Collaborate closely with engineers across deep learning frameworks, libraries, kernels, compilers, and GPU architecture teams at NVIDIA. * Contribute to open-source communities and ecosystem integrations where relevant, including projects such as FlashInfer, vLLM, and SGLang. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Just-in-time Compilation in JVM](https://www.wearedevelopers.com/videos/240-just-in-time-compilation-in-jvm) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python)