> Markdown version of [/jobs/ext/2028710-principal-software-development-engineer-ai-ml](https://www.wearedevelopers.com/jobs/ext/2028710-principal-software-development-engineer-ai-ml). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Development Engineer - AI/ML - **Company:** Advanced Micro Devices, Inc. - **Location:** San Jose, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Compilers, Profiling, Nvidia CUDA, Software Debugging, Linux, Object-Oriented Software Development, Performance Tuning, Tensorflow, Software Engineering, Graphics Processing Unit (GPU), Gpu Programming, Git, ONNX (Open Neural Network Exchange) Format - **Published:** August 11, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/principal-software-development-engineer-ai-ml-san-jose-ca-usa-58901326 ## About the Role training, fine-tuning, RL, and inference on AMD GPUs * Design and implement graph-, kernel-, and runtime-level optimizations yielding measurable gains * Analyze workload characteristics to identify compute, memory, and communication improvements * Drive performance characterization and benchmarking pre-silicon and post-silicon Tasks * Strong proficiency in C++ and object-oriented software development * Hands-on GPU programming experience (HIP, CUDA, Triton, or equivalent) * Experience with MLIR, LLVM, or related compiler infrastructures * Solid understanding of ML frameworks, model execution pipelines, and AI workload optimization * Experience profiling and optimizing performance-critical software * Proficiency with Linux, Git, debugging tools, and performance profilers * Strong analytical, communication, and problem-solving skills Key requirements * ## Description Experteer Overview In this role, you will optimize AI workloads on AMD GPUs through advanced compiler work and performance engineering. You'll bridge architecture, compilers, and ML frameworks to enhance ROCm and ONNX Runtime execution. You will collaborate across architecture, kernel, runtime, and framework teams to deliver scalable AI solutions. This position offers the chance to shape AI performance across current and future GPU platforms. Join a mission-driven team pushing the boundaries of AI on AMD hardware. Compensation / Benefits * Collaborate with architecture specialists to shape future products * Design, develop, and optimize compiler transformations and operator lowerings within MLIR, LLVM, and ROCm * Create scalable solutions for portability and performance across AMD GPUs * Work with compiler and runtime teams on advanced AI workload optimizations * Implement, validate, and maintain ONNX operators and microbenchmark suites * Lead performance optimization efforts for AI training, fine-tuning, RL, and inference on AMD GPUs * Design and implement graph-, kernel-, and runtime-level optimizations yielding measurable gains * Analyze workload characteristics to identify compute, memory, and communication improvements * Drive performance characterization and benchmarking pre-silicon and post-silicon Tasks * Strong proficiency in C++ and object-oriented software development * Hands-on GPU programming experience (HIP, CUDA, Triton, or equivalent) * Experience with MLIR, LLVM, or related compiler infrastructures * Solid understanding of ML frameworks, model execution pipelines, and AI workload optimization * Experience profiling and optimizing performance-critical software * Proficiency with Linux, Git, debugging tools, and performance profilers * Strong analytical, communication, and problem-solving skills Key requirements * ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Just-in-time Compilation in JVM](https://www.wearedevelopers.com/videos/240-just-in-time-compilation-in-jvm) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs](https://www.wearedevelopers.com/magazine/557-dev-digest-157-cuda-in-python-gemini-code-assist-and-back-dooring-llms) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Transforming Software Development: The Role of AI and Developer Tools](https://www.wearedevelopers.com/magazine/527-transforming-software-development-the-role-of-ai-and-developer-tools)