> Markdown version of [/jobs/ext/2058196-software-engineer-gpu-ai-ml](https://www.wearedevelopers.com/jobs/ext/2058196-software-engineer-gpu-ai-ml). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer- GPU/AI/ML - **Company:** Advanced Micro Devices, Inc. - **Location:** Santa Clara, CA, United States - **Salary:** $204,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Profiling, Nvidia CUDA, Computer Engineering, Distributed Computing Environment, Distributed Systems, General-Purpose Computing on Graphics Processing Units, Hardware Design, Tensorflow, Software Engineering, SystemVerilog, Verilog, Graphics Processing Unit (GPU), Pytorch, Delivery Pipeline, Large Language Models, Deep Learning, Information Technology - **Published:** August 14, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/87575403/1 ## About the Role We need someone who can go deep in the areas below and collaborate effectively. SYSTEMS & GPU PERFORMANCE (KERNEL ENGINEERING) * Expert-level modern C++ and design of large, performance-critical systems. * Strong grasp of GPU architecture, memory hierarchy, and kernel optimization (HIP/CUDA). * Hands-on delivery on large-scale C++/HIP/CUDA codebases, such as ROCm (rocBLAS, hipDNN, Composable Kernel, AITemplate), the CUDA ecosystem (cuBLAS, cuDNN, CUTLASS, Thrust, CUB, NCCL), and ML framework cores such as PyTorch, TensorFlow, or JAX (C++/HIP/CUDA paths). * Comfort diagnosing bottlenecks with profilers (for example, ROCm Profiler and Nsight) in multi-GPU, distributed settings. AI POST-TRAINING & LLM SYSTEMS * Deep understanding of transformers, attention, and the full model lifecycle. * Hands-on work in alignment and post-training-for example, SFT, RLHF, and GRPO. * Awareness of current LLM trends, including MoE, quantization, speculative decoding, and agentic systems. * Experience optimizing post-training and inference pipelines at scale. PREFERRED BACKGROUND * Substantial professional experience in software development within performance-critical environments. * Extensive HIP/CUDA experience optimizing deep learning and OSS LLM inference/training kernels and operators. * Strong technical ownership and a track record of shipping complex systems. * Clear communication and influence across teams. * Plus: Deep familiarity with the AMD ROCm/HIP ecosystem. * Plus: Working knowledge of RTL design and Verilog/SystemVerilog for hardware-software co-design., * Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience. * Master's preferred; PhD a plus. * Publications in AI/ML, GPU computing, or systems optimization are valued. ## Description We're looking for a senior software engineer who combines deep systems performance work with modern AI-someone who can shape software from GPU kernels through distributed training and inference. You'll join a core team of specialists working on the latest AMD hardware and software. Your work will directly influence the ROCm ecosystem and how foundation models and agentic systems perform on AMD GPUs. The challenge: Help train and run AI systems that make AI itself more efficient on GPUs-tuning stacks, kernels, and workflows in ways that can materially shift what's possible on our hardware. This is a high-impact, hands-on role. You'll own hard technical problems, influence direction across teams, and mentor others as we scale AMD's AI software strategy., * Own the AI software stack: Establish best practices and drive performance from low-level GPU kernels to large-scale distributed systems. Use modern LLMs and agent-based tooling where it accelerates development and tuning of the ROCm ecosystem. * Accelerate foundation models and agents: Improve training, post-training, and inference for LLMs and autonomous AI workloads so AMD is the default platform for the most demanding use cases. * Co-design hardware and software: Partner on the full lifecycle-from GPU architecture input to software for new accelerators-and engage with the broader AI community to keep AMD at the forefront., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Enhancing Workload Security in Kubernetes](https://www.wearedevelopers.com/videos/356-enhancing-workload-security-in-kubernetes) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs](https://www.wearedevelopers.com/magazine/557-dev-digest-157-cuda-in-python-gemini-code-assist-and-back-dooring-llms) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)