> Markdown version of [/jobs/ext/2050048-frontier-ai-workloads-performance-and-scalability-engineer](https://www.wearedevelopers.com/jobs/ext/2050048-frontier-ai-workloads-performance-and-scalability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Frontier AI Workloads - Performance and Scalability Engineer - **Company:** Advanced Micro Devices, Inc. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $226,400.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Computer Engineering, Microarchitecture, Linux, Distributed Systems, General-Purpose Computing on Graphics Processing Units, High-Level Architecture, Python (Programming Language), Open Source Technology, OpenCL, Remote Direct Memory Access, Software Engineering, Systems Architecture, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Information Technology, Free and Open-Source Software - **Published:** August 14, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/86406449/1 ## About the Role The ideal candidate is a deep technical expert with a track record of solving industry-hard problems at the intersection of GPU architecture, AI systems, and high-performance software. You understand the full stack from hardware micro-architecture to model architecture, inference paradigms, and system-level design. You lead through technical depth, influence, and by example, staying hands-on while setting direction for your team. If you want to shape how the world runs AI on AMD hardware, this role is for you., * Deep software development experience in GPU computing, HPC, or AI systems * Deep understanding of GPU micro-architecture, memory hierarchy, instruction scheduling, and performance tradeoffs * Deep understanding of end-to-end AI systems: model architectures, inference paradigms, and system/rack-level design * Understanding of multi-GPU communication: scale-up (NVLink, xGMI, Infinity Fabric) and scale-out (RDMA, RCCL/NCCL) topologies and performance characteristics * Experience designing and optimizing across the full stack: from low-level GPU kernels to frameworks and distributed serving systems * Strong background in performance engineering, including profiling, roofline analysis, and bottleneck diagnosis at scale * Experience with one or more of: HIP, CUDA, OpenCL, Triton/Gluon, CUTLASS, CK * Experience with GPU compiler toolchains (e.g., LLVM) and intermediate representations (e.g., MLIR, LLVM IR, Triton IR) is a plus * Hands-on experience contributing to or architecting major open-source AI frameworks (e.g., vLLM, SGLang, xDiT, Megatron LM, PyTorch) * Strong proficiency in C++ (C++17 or later) and Python * Experience leading small technical teams while remaining a hands-on contributor * Track record of influencing technical direction across teams and organizations * Strong Linux systems knowledge * Excellent written and verbal English communication skills * Published research or significant open-source contributions in GPU computing, HPC, or AI systems is a plus PREFERRED ACADEMIC CREDENTIALS: * Master's or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent. PhD strongly preferred. ## Description AMD is looking for a Principal Engineer to serve as a hands-on technical team lead driving the performance and scalability of frontier AI workloads on AMD GPUs, including large language models, mixture-of-experts architectures, and diffusion models. You will lead a team of engineers, define the long-term technical vision, make critical architecture decisions, and tackle the hardest performance challenges across the stack from GPU kernels and to serving frameworks and distributed systems., * Lead a small team of engineers: set technical direction, prioritize work, and ensure delivery while remaining deeply hands-on * Define and drive the long-term technical strategy for AI workload performance on AMD GPUs * Own the most complex cross-stack performance challenges, from kernel optimization to framework-level architecture decisions * Lead the design and implementation of novel GPU kernels, compiler optimizations, and framework features * Establish performance methodology and roofline analysis practices that set the standard for the team * Influence upstream roadmaps in major open-source AI frameworks (e.g., vLLM, SGLang, PyTorch) * Drive architecture decisions for emerging inference paradigms (e.g., prefill-decode disaggregation, speculative decoding, distributed serving) * Identify and close fundamental performance gaps between AMD and competitor platforms * Serve as a technical authority across the organization, advising leadership on technical direction and feasibility * Mentor engineers and raise the technical bar across the broader engineering organization * Represent AMD externally through publications, conference talks, and open-source contributions, AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)