> Markdown version of [/jobs/ext/1890911-fellow-software-engineer-ai-performance-reliability](https://www.wearedevelopers.com/jobs/ext/1890911-fellow-software-engineer-ai-performance-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Fellow Software Engineer - AI Performance & Reliability - **Company:** Advanced Micro Devices, Inc. - **Location:** San Jose, CA, United States - **Salary:** $268,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, C++ (Programming Language), Program Optimization, Profiling, Nvidia CUDA, Software Debugging, Distributed Systems, Monitoring of Systems, Python (Programming Language), Machine Learning, Performance Tuning, Recommender Systems, Tensorflow, Software Engineering, AI Infrastructure, Pytorch, Large Language Models, Parallel Computation, Reliability of Systems, Information Technology, Low Latency, Machine Learning Operations, Service Stack, Programming Languages - **Published:** August 2, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17784482?backUrl=%2Fcareer%2F17784482%2FFellow-Software-Engineer-Ai-Performance-Reliability-California-San-Jose ## About the Role * Strong software engineering skills and experience building production-quality systems. * Experience working with AI infrastructure for model training, inference, or both. * Demonstrated experience profiling and optimizing machine learning models or AI workloads. * Strong foundations in computer architecture, including processors, memory hierarchies, parallelism, and performance tradeoffs. * Solid understanding of systems performance concepts such as latency, throughput, memory bandwidth, utilization, and distributed communication. * Proficiency in languages such as Python, C++, or similar systems-oriented programming languages. * Experience with machine learning frameworks such as PyTorch, TensorFlow, or JAX. * Strong debugging and analytical skills, with the ability to investigate problems across multiple layers of the technology stack. * Clear written and verbal communication skills. * A customer-focused mindset and willingness to work directly with customers through technical evaluations, deployments, troubleshooting, and ongoing support. PREFERRED EXPERIENCE: * Experience optimizing large language models, diffusion models, or recommendation models. * Experience with GPU, accelerator, or distributed computing environments. * Familiarity with technologies such as ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL, or similar performance-oriented tools and runtimes. * Experience with distributed training, model serving, quantization, compilation, kernel optimization, or memory optimization. * Experience operating AI systems in production environments. * Prior experience in solutions engineering, field engineering, developer relations, or another customer-facing technical role. * Experience designing benchmarks and conducting systematic performance analysis. ACADEMIC CREDENTIALS: * A PhD (or a master's degree with equivalent experience) in artificial intelligence, machine learning, computer science, or a related field. ## Description We are looking for a strong, Principal or Fellow level software engineer to join our AI Infrastructure team. You will work on improving the performance, efficiency, and reliability of AI workloads across both model training and inference. Our team supports a broad range of machine learning systems, including large language models, diffusion models, and recommendation models. You will collaborate closely with customers and internal engineering teams to understand performance bottlenecks, optimize workloads, and ensure that models run reliably at scale. This role is a strong fit for an engineer who enjoys working across the AI software and hardware stack, solving technically challenging performance problems, and partnering directly with customers to make them successful. You will help customers achieve meaningful improvements in model performance and system reliability. You will identify difficult bottlenecks, develop reusable solutions, and help shape the infrastructure and product capabilities needed to run demanding AI workloads efficiently at scale. THE PERSON: * Profile and optimize AI model training and inference workloads. * Improve model throughput, latency, memory efficiency, scalability, and reliability. * Identify bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware. * Optimize workloads involving large language models, diffusion models, recommendation systems, and other modern machine learning architectures. * Develop performance tooling, benchmarks, automation, and observability systems. * Investigate and resolve complex production issues affecting AI workloads. * Collaborate with customers to understand their technical requirements, reproduce issues, and recommend effective solutions. * Translate customer feedback into product and infrastructure improvements. * Work closely with machine learning engineers, systems engineers, hardware teams, and product teams. * Document performance findings, technical recommendations, and best practices., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. ## Related Videos - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)