Systems Design Engineer (AI, Software)

Advanced Micro Devices, Inc.
San Jose, CA, United States
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Adobe Analytics Board Bringup Artificial Intelligence C++ (Programming Language) Profiling Computer Programming Software Debugging Data Flow Control Python (Programming Language) Linux Kernel Machine Learning Performance Tuning
+8 more
Software Engineering Graphics Processing Unit (GPU) Pytorch Parallel Computation Linux Development ONNX (Open Neural Network Exchange) Format Machine Learning Operations Software Version Control

Job description

Experteer Overview In this role you develop and optimize ML workloads on AMD AI accelerators, bridging hardware and software. You design high-performance ML operator kernels and dataflow libraries, and work with cross-functional teams to bring AI tech from concept to production. You gain full-stack visibility from kernel development to silicon bring-up, contributing to industry-leading AI inference performance. This is a chance to impact products deployed in millions of devices worldwide and help shape accelerator technology. Compensation / Benefits * Develop and optimize ML operator kernels and dataflow libraries for AMD AI accelerators * Profile workloads to identify bottlenecks and drive system-level optimizations * Enable and validate ML models within production inference frameworks and runtimes * Collaborate with compiler, runtime, architecture, and silicon teams to deliver high-performance AI solutions * Debug and resolve issues across kernel implementation, runtime integration, model accuracy, and hardware bring-up * Contribute to hardware-software co-design by evaluating architectural tradeoffs * Drive innovation in performance methodologies, benchmarking, tooling, and AI system optimization Tasks * Strong software development experience using C/C++ and Python * Experience with parallel programming and performance optimization * Knowledge of ML inference workloads and common operators (e.g., GEMM, convolution, attention, softmax) * Familiarity with AI frameworks/runtimes (PyTorch, ONNX Runtime, ROCm) * Understanding of computer architecture, memory hierarchies, and accelerator programming models * Experience developing software for GPUs, NPUs, or AI accelerators * Experience with Linux development, debugging, profiling, and source control tools * Familiarity with MLIR, LLVM, compiler technologies, or related stacks * Exposure to quantization techniques (INT8, FP8, FP16, BF16) * Knowledge of dataflow architectures, systolic arrays, or custom accelerators * Publications, patents, or demonstrated contributions in ML systems or computer architecture Key requirements *

Requirements

model accuracy, and hardware bring-up * Contribute to hardware-software co-design by evaluating architectural tradeoffs * Drive innovation in performance methodologies, benchmarking, tooling, and AI system optimization Tasks * Strong software development experience using C/C++ and Python * Experience with parallel programming and performance optimization * Knowledge of ML inference workloads and common operators (e.g., GEMM, convolution, attention, softmax) * Familiarity with AI frameworks/runtimes (PyTorch, ONNX Runtime, ROCm) * Understanding of computer architecture, memory hierarchies, and accelerator programming models * Experience developing software for GPUs, NPUs, or AI accelerators * Experience with Linux development, debugging, profiling, and source control tools * Familiarity with MLIR, LLVM, compiler technologies, or related stacks * Exposure to quantization techniques (INT8, FP8, FP16, BF16) * Knowledge of dataflow architectures, systolic arrays, or custom accelerators * aa programming patents, or demonstrated contributions in ML systems or computer architecture Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:20 min

Utilizing AI and hardware acceleration for application code optimization

Stephan Gillich Stephan Gillich · WWC 2024

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · WWC Europe 2026

1:15 min

Overcoming the challenges of modifying Linux kernel code

Ayesha Kaleem · WWC 2023

3:18 min

Hardware architectures tailored for specific artificial intelligence computations

Stephan Gillich Stephan Gillich · WWC 2024

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

Videos

See all

Related articles

See all