> Markdown version of [/jobs/ext/1788403-software-engineer-kernels-cuda-c](https://www.wearedevelopers.com/jobs/ext/1788403-software-engineer-kernels-cuda-c). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - Kernels/CUDA (C++) - **Company:** SPACEXAI LLC - **Location:** Palo Alto, CA, United States - **Salary:** $180,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Computer Clusters, Nvidia CUDA, Software Debugging, File Systems, Linux Kernel, Supercomputing, System Programming, Syntactically Awesome Style Sheets (SASS) - **Published:** July 14, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3318672948&tx=YT1212TYD&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates., * Deep low-level systems programming (C/C++/PTX/SASS) * Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale * Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling) * Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms) * Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios ## Description We are building one of the world's largest AI supercomputers from the ground up. As part of the Compute Infrastructure team, you will own both the raw GPU supercomputer and the platform layer that runs on top of it. You will work across the full stack - from low-level GPU kernel optimizations and Linux kernel internals to massive-scale orchestration and virtualization - to make training and inference at xAI as fast, reliable, and scalable as possible. This is a broad, high-impact role that combines hardcore supercompute and compute infrastructure work. Your contributions will directly accelerate Grok's training speed and overall AI progress. RESPONSIBILITIES * Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads * Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance * Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operation * Collaborate closely with AI research teams to deliver production-grade performance and scalability ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Possibilities with Web Capabilities](https://www.wearedevelopers.com/videos/1575-possibilities-with-web-capabilities) - [The weekly developer show: Boosting Python with CUDA, CSS Updates & Navigating New Tech Stacks](https://www.wearedevelopers.com/videos/1293-the-weekly-developer-show-boosting-python-with-cuda-css-updates-navigating-new-tech-stacks) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) ## Related Articles - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)