> Markdown version of [/jobs/ext/2122984-gpu-software-engineer-cuda](https://www.wearedevelopers.com/jobs/ext/2122984-gpu-software-engineer-cuda). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # GPU Software Engineer (CUDA) - **Company:** Bright Vision Technologies - **Location:** Novi, MI, United States (Remote available) - **Experience:** Expert - **Salary:** $80,000.0 - $107,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Profiling, Code Review, Nvidia CUDA, Computer Programming, Computer Engineering, Tensorflow, Scientific Computating, Software Engineering, System Programming, Data Processing, Graphics Processing Unit (GPU), High Performance Computing, Gpu Programming, Information Technology, Free and Open-Source Software, TensorRT - **Published:** August 19, 2026 - **Apply:** https://www.careerjet.com/jobad/us98b31cfc578180bbc73dd48922477326 ## About the Role * Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field. * Six or more years of experience in GPU programming and performance engineering. * Deep expertise in CUDA C/C++ and GPU programming models. * Strong understanding of modern GPU architectures, memory hierarchies, and execution models. * Hands-on experience profiling and optimizing GPU workloads in production. * Familiarity with NCCL, MPI, and high-performance interconnect technologies. * Experience integrating custom kernels into ML frameworks. * Strong C++ skills and familiarity with modern systems programming practices. * Solid grounding in linear algebra and numerical methods. * Strong communication and collaboration skills with research and engineering teams. Preferred Qualifications * Experience with Triton, CUTLASS, or other GPU kernel authoring frameworks. * Familiarity with TensorRT, FasterTransformer, or vLLM internals. * Exposure to compiler infrastructure such as LLVM or MLIR. * Open-source contributions to GPU or ML performance libraries. * Experience with large-scale distributed training infrastructure. ## Description We are seeking a GPU Software Engineer (CUDA) with deep expertise in CUDA programming, GPU architecture, and high-performance computing to design and optimize compute-intensive workloads on modern accelerator hardware. This role focuses on extracting maximum performance from GPU platforms for AI training, inference, scientific computing, and high-throughput data processing workloads. The ideal candidate combines low-level systems mastery with strong software engineering practices, and has a track record of delivering measurable performance improvements on production GPU systems. In this role you will work closely with cross-functional partners - product, design, engineering, operations, and business stakeholders - to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs](https://www.wearedevelopers.com/magazine/557-dev-digest-157-cuda-in-python-gemini-code-assist-and-back-dooring-llms) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers)