Systems Design Engineer (AI, Software)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+8 more
Job description
Experteer Overview In this role you develop and optimize ML workloads on AMD AI accelerators, bridging hardware and software. You design high-performance ML operator kernels and dataflow libraries, and work with cross-functional teams to bring AI tech from concept to production. You gain full-stack visibility from kernel development to silicon bring-up, contributing to industry-leading AI inference performance. This is a chance to impact products deployed in millions of devices worldwide and help shape accelerator technology. Compensation / Benefits * Develop and optimize ML operator kernels and dataflow libraries for AMD AI accelerators * Profile workloads to identify bottlenecks and drive system-level optimizations * Enable and validate ML models within production inference frameworks and runtimes * Collaborate with compiler, runtime, architecture, and silicon teams to deliver high-performance AI solutions * Debug and resolve issues across kernel implementation, runtime integration, model accuracy, and hardware bring-up * Contribute to hardware-software co-design by evaluating architectural tradeoffs * Drive innovation in performance methodologies, benchmarking, tooling, and AI system optimization Tasks * Strong software development experience using C/C++ and Python * Experience with parallel programming and performance optimization * Knowledge of ML inference workloads and common operators (e.g., GEMM, convolution, attention, softmax) * Familiarity with AI frameworks/runtimes (PyTorch, ONNX Runtime, ROCm) * Understanding of computer architecture, memory hierarchies, and accelerator programming models * Experience developing software for GPUs, NPUs, or AI accelerators * Experience with Linux development, debugging, profiling, and source control tools * Familiarity with MLIR, LLVM, compiler technologies, or related stacks * Exposure to quantization techniques (INT8, FP8, FP16, BF16) * Knowledge of dataflow architectures, systolic arrays, or custom accelerators * Publications, patents, or demonstrated contributions in ML systems or computer architecture Key requirements *
Requirements
model accuracy, and hardware bring-up * Contribute to hardware-software co-design by evaluating architectural tradeoffs * Drive innovation in performance methodologies, benchmarking, tooling, and AI system optimization Tasks * Strong software development experience using C/C++ and Python * Experience with parallel programming and performance optimization * Knowledge of ML inference workloads and common operators (e.g., GEMM, convolution, attention, softmax) * Familiarity with AI frameworks/runtimes (PyTorch, ONNX Runtime, ROCm) * Understanding of computer architecture, memory hierarchies, and accelerator programming models * Experience developing software for GPUs, NPUs, or AI accelerators * Experience with Linux development, debugging, profiling, and source control tools * Familiarity with MLIR, LLVM, compiler technologies, or related stacks * Exposure to quantization techniques (INT8, FP8, FP16, BF16) * Knowledge of dataflow architectures, systolic arrays, or custom accelerators * aa programming patents, or demonstrated contributions in ML systems or computer architecture Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Stephan Gillich - Bringing AI Everywhere
Highest Paying Tech Companies for Developers
Dev Digest 120 - Apple and peers