Software Engineer, CUDA Deep Learning Systems
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview In this role you will advance deep learning workloads by optimizing CUDA-based systems for cutting-edge AI models. You will work with a cross-functional team to prototype high-performance kernels and distributed pipelines that scale from a single node to clusters. The job blends research and practical implementation, aiming to maximize accelerator utilization and memory bandwidth across training and inference. This is a chance to shape next-generation AI systems on modern GPUs and contribute to open-source and internal tooling. You will join a highly technical, research-oriented group tackling uncharted optimization and architecture challenges. Compensation / Benefits * Explore and prototype system optimizations at the intersection of high-level DL frameworks and CUDA * Architect and optimize distributed computing systems from single-node to cluster-scale * Design, implement, and optimize custom high-performance CUDA kernels * Analyze hardware-software interactions to identify bottlenecks in training and inference * Collaborate with AI researchers, HW/SW architects, kernel and compiler experts, and CUDA drivers * Develop exploratory tools and runtime systems to profile and accelerate new DL paradigms * Write clean, maintainable code to enable prototypes to transition to open-source releases or products Tasks * BS/MS/PhD in CS, CE, EE, or related field (or equivalent experience) * 2+ years of relevant industry or academic experience * Strong proficiency in C++ and Python * Solid fundamentals in Deep Learning with a focus on transformers * Strong understanding of distributed computing, multi-node scaling, and performance challenges in cluster environments * Proven experience in systems programming, computer architecture, and low-level performance optimization * Hands-on experience with CUDA programming, kernel optimization, and workload profiling Key requirements * equity * benefits * remote/hybrid options * competitive compensation
Requirements
_ to identify bottlenecks in training and inference * Collaborate with AI researchers, HW/SW architects, kernel and compiler experts, and CUDA drivers * Develop exploratory tools and runtime systems to profile and accelerate new DL paradigms * Write clean, maintainable code to enable prototypes to transition to open-source releases or products Tasks * BS/MS/PhD in CS, CE, EE, or related field (or equivalent experience) * 2+ years of relevant industry or academic experience * Strong proficiency in C++ and Python * Solid fundamentals in Deep Learning with a focus on transformers * Strong understanding of distributed computing, multi-node scaling, and performance challenges in cluster environments * Proven experience in systems programming, computer architecture, and low-level performance optimization * Hands-on experience with CUDA programming, kernel optimization, and workload profiling Key requirements * equity * benefits * remote/hybrid options * competitive compensation
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
How to Become an AI Engineer
Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs
MLOps And AI Driven Development