Software Engineer, CUDA Deep Learning Systems
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview In this role you will advance deep learning workloads by optimizing CUDA-based systems for cutting-edge AI models. You will work with a cross-functional team to prototype high-performance kernels and distributed pipelines that scale from a single node to clusters. The job blends research and practical implementation, aiming to maximize accelerator utilization and memory bandwidth across training and inference. This is a chance to shape next-generation AI systems on modern GPUs and contribute to open-source and internal tooling. You will join a highly technical, research-oriented group tackling uncharted optimization and architecture challenges. Compensation / Benefits * Explore and prototype system optimizations at the intersection of high-level DL frameworks and CUDA * Architect and optimize distributed computing systems from single-node to cluster-scale * Design, implement, and optimize custom high-performance CUDA kernels * Analyze hardware-software interactions to identify bottlenecks in training and inference * Collaborate with AI researchers, HW/SW architects, kernel and compiler experts, and CUDA drivers * Develop exploratory tools and runtime systems to profile and accelerate new DL paradigms * Write clean, maintainable code to enable prototypes to transition to open-source releases or products Tasks * BS/MS/PhD in CS, CE, EE, or related field (or equivalent experience) * 2+ years of relevant industry or academic experience * Strong proficiency in C++ and Python * Solid fundamentals in Deep Learning with a focus on transformers * Strong understanding of distributed computing, multi-node scaling, and performance challenges in cluster environments * Proven experience in systems programming, computer architecture, and low-level performance optimization * Hands-on experience with CUDA programming, kernel optimization, and workload profiling Key requirements * equity * benefits * remote/hybrid options * competitive compensation
Requirements
_ to identify bottlenecks in training and inference * Collaborate with AI researchers, HW/SW architects, kernel and compiler experts, and CUDA drivers * Develop exploratory tools and runtime systems to profile and accelerate new DL paradigms * Write clean, maintainable code to enable prototypes to transition to open-source releases or products Tasks * BS/MS/PhD in CS, CE, EE, or related field (or equivalent experience) * 2+ years of relevant industry or academic experience * Strong proficiency in C++ and Python * Solid fundamentals in Deep Learning with a focus on transformers * Strong understanding of distributed computing, multi-node scaling, and performance challenges in cluster environments * Proven experience in systems programming, computer architecture, and low-level performance optimization * Hands-on experience with CUDA programming, kernel optimization, and workload profiling Key requirements * equity * benefits * remote/hybrid options * competitive compensation
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role β technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
How to Become an AI Engineer
Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs
MLOps And AI Driven Development