Systems Design Engineer (AI, Software)

Advanced Micro Devices, Inc.
San Jose, CA, United States
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Board Bringup Artificial Intelligence C++ (Programming Language) Profiling Computer Programming Computer Engineering Software Debugging Data Flow Control Python (Programming Language) Linux Kernel Linux System Administration Machine Learning
+13 more
Performance Tuning Software Engineering Multithreading Graphics Processing Unit (GPU) High Performance Computing Pytorch Model Validation Hardware Testing Parallel Computation Information Technology ONNX (Open Neural Network Exchange) Format Machine Learning Operations Software Version Control

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.

AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger - technology that moves the world forward.

THE ROLE

AMD is seeking an AI Systems Engineer to help develop and optimize machine learning workloads on next-generation AMD AI accelerators. In this role, you will work at the intersection of hardware and software, designing high-performance ML operator kernels, optimizing dataflow pipelines, and enabling industry-leading AI inference performance across AMD NPU and GPU platforms.

You will collaborate closely with compiler, runtime, silicon, and architecture teams while helping bring cutting-edge AI technologies from concept to production. This role offers full-stack visibility from kernel development and model optimization through hardware validation and silicon bring-up. If you are passionate about AI systems, accelerator architectures, and solving complex performance challenges, this is an opportunity to make a significant impact on products deployed in millions of devices worldwide. THE PERSON

The ideal candidate is a systems-minded engineer who enjoys tackling complex performance and optimization challenges at the hardware-software boundary. You are naturally curious, thrive in collaborative environments, and are comfortable working across multiple technical domains to debug, analyze, and improve system behavior.

You have a strong foundation in computer architecture and machine learning systems, enjoy working with cross-functional teams, and can translate technical insights into scalable solutions that improve performance, functionality, and product quality. KEY RESPONSIBILITIES

  • Develop and optimize machine learning operator kernels and dataflow libraries for AMD AI accelerators.
  • Profile workloads, identify performance bottlenecks, and drive software and system-level optimizations.
  • Enable and validate ML models within production inference frameworks and runtime environments.
  • Collaborate with compiler, runtime, architecture, and silicon teams to deliver high-performance AI solutions.
  • Debug and resolve issues spanning kernel implementation, runtime integration, model accuracy, and hardware bring-up.
  • Contribute to hardware-software co-design efforts by evaluating architectural tradeoffs and influencing future accelerator technologies.
  • Drive innovation in performance methodologies, benchmarking, tooling, and AI system optimization.

PREFERRED EXPERIENCE

  • Strong software development experience using C/C++ and Python.
  • Experience with parallel programming, multithreaded applications, and performance optimization.
  • Knowledge of machine learning inference workloads and common operators such as GEMM, convolution, attention, and softmax.
  • Familiarity with AI frameworks and runtimes such as PyTorch, ONNX Runtime, ROCm, or similar technologies.
  • Understanding of computer architecture, memory hierarchies, cache behavior, and accelerator programming models.
  • Experience developing software for GPUs, NPUs, AI accelerators, or other high-performance computing platforms.
  • Experience using development, debugging, profiling, and source control tools in Linux environments.
  • Familiarity with MLIR, LLVM, compiler technologies, or related software stacks.
  • Exposure to quantization techniques, including INT8, FP8, FP16, or BF16 optimization.
  • Knowledge of dataflow architectures, systolic arrays, or custom accelerator designs.
  • Publications, patents, or demonstrated technical contributions in machine learning systems, computer architecture, or related fields.

ACADEMIC CREDENTIALS

  • Master’s or PhD in Computer Engineering, Electrical Engineering, Computer Science, or a related technical field preferred.

LOCATION

San Jose, CA

This role is not eligible for visa sponsorship.

LI-DR2

LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Requirements

  • Strong software development experience using C/C++ and Python.
  • Experience with parallel programming, multithreaded applications, and performance optimization.
  • Knowledge of machine learning inference workloads and common operators such as GEMM, convolution, attention, and softmax.
  • Familiarity with AI frameworks and runtimes such as PyTorch, ONNX Runtime, ROCm, or similar technologies.
  • Understanding of computer architecture, memory hierarchies, cache behavior, and accelerator programming models.
  • Experience developing software for GPUs, NPUs, AI accelerators, or other high-performance computing platforms.
  • Experience using development, debugging, profiling, and source control tools in Linux environments.
  • Familiarity with MLIR, LLVM, compiler technologies, or related software stacks.
  • Exposure to quantization techniques, including INT8, FP8, FP16, or BF16 optimization.
  • Knowledge of dataflow architectures, systolic arrays, or custom accelerator designs.
  • Publications, patents, or demonstrated technical contributions in machine learning systems, computer architecture, or related fields.

ACADEMIC CREDENTIALS

  • Master’s or PhD in Computer Engineering, Electrical Engineering, Computer Science, or a related technical field preferred.

About the company

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.

AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger - technology that moves the world forward.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerarc.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:54 min

Leveraging chip sets and software for artificial intelligence efficiency

Ankit Patel Ankit Patel ¡ WWC 2025

2:35 min

Preventing remote code execution in PyTorch models

Balåzs Kiss ¡ WWC 2023

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar ¡ WWC Europe 2026

1:15 min

Overcoming the challenges of modifying Linux kernel code

Ayesha Kaleem ¡ WWC 2023

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa ¡ LIVE

2:42 min

Dissecting artificial intelligence layers from compute to applications

Christian Nagel Christian Nagel +3 ¡ WWC Europe 2026

Videos

See all

Related articles

See all