Principal Software Development Engineer - AI/ML
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
Experteer Overview In this role, you will optimize AI workloads on AMD GPUs through advanced compiler work and performance engineering. You’ll bridge architecture, compilers, and ML frameworks to enhance ROCm and ONNX Runtime execution. You will collaborate across architecture, kernel, runtime, and framework teams to deliver scalable AI solutions. This position offers the chance to shape AI performance across current and future GPU platforms. Join a mission-driven team pushing the boundaries of AI on AMD hardware. Compensation / Benefits * Collaborate with architecture specialists to shape future products * Design, develop, and optimize compiler transformations and operator lowerings within MLIR, LLVM, and ROCm * Create scalable solutions for portability and performance across AMD GPUs * Work with compiler and runtime teams on advanced AI workload optimizations * Implement, validate, and maintain ONNX operators and microbenchmark suites * Lead performance optimization efforts for AI training, fine-tuning, RL, and inference on AMD GPUs * Design and implement graph-, kernel-, and runtime-level optimizations yielding measurable gains * Analyze workload characteristics to identify compute, memory, and communication improvements * Drive performance characterization and benchmarking pre-silicon and post-silicon Tasks * Strong proficiency in C++ and object-oriented software development * Hands-on GPU programming experience (HIP, CUDA, Triton, or equivalent) * Experience with MLIR, LLVM, or related compiler infrastructures * Solid understanding of ML frameworks, model execution pipelines, and AI workload optimization * Experience profiling and optimizing performance-critical software * Proficiency with Linux, Git, debugging tools, and performance profilers * Strong analytical, communication, and problem-solving skills Key requirements *
Requirements
training, fine-tuning, RL, and inference on AMD GPUs * Design and implement graph-, kernel-, and runtime-level optimizations yielding measurable gains * Analyze workload characteristics to identify compute, memory, and communication improvements * Drive performance characterization and benchmarking pre-silicon and post-silicon Tasks * Strong proficiency in C++ and object-oriented software development * Hands-on GPU programming experience (HIP, CUDA, Triton, or equivalent) * Experience with MLIR, LLVM, or related compiler infrastructures * Solid understanding of ML frameworks, model execution pipelines, and AI workload optimization * Experience profiling and optimizing performance-critical software * Proficiency with Linux, Git, debugging tools, and performance profilers * Strong analytical, communication, and problem-solving skills Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs
Stephan Gillich - Bringing AI Everywhere
Dev Digest 120 - Apple and peers