Principal System Software Architect, AI/GPU Platforms

Advanced Micro Devices, Inc.
Austin, TX, United States
16 days ago
Apply on diversityjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$204,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Computing Platforms Nvidia CUDA Computer Engineering Data Centers Linux Device Drivers Distributed Computing Environment Memory Management Ethernet Fault Tolerance Firmware
+16 more
InfiniBand Machine Learning Node.Js Open Source Technology OpenCL Performance Tuning Remote Direct Memory Access Tensorflow Software Engineering System Software Graphics Processing Unit (GPU) Computer Network Technologies Pytorch Deep Learning Information Technology ONNX (Open Neural Network Exchange) Format

Job description

You will join the system software architecture team behind AMD Instinct accelerators, the GPUs powering some of the world’s largest AI and HPC deployments. This role is focused on next-generation, rack-scale AI platforms in the MI400 class spanning the GPU, the node, and the scale-up/scale-out fabric that binds thousands of accelerators into a single training and inference system. THE PERSON

As a Systems Software Architect, you sit at the intersection of silicon, firmware, driver, runtime, and framework. You define how the software stack exposes and orchestrates the hardware so that AMD’s largest customers can extract maximum performance, reliability, and utilization from their infrastructure. KEY RESPONSIBILITIES

  • Own the end-to-end system software architecture for one or more MI400-class subsystems for example GPU memory management, scheduling and queuing, RAS and serviceability, virtualization/partitioning (SR-IOV), or the scale-up/scale-out interconnect software model.
  • Drive architecture across the stack: kernel-mode driver (amdgpu/KFD), user-mode runtime (ROCr/HSA), firmware interfaces, and the ROCm software platform, ensuring the layers compose cleanly and perform.
  • Partner with silicon and SoC architects during pre-silicon definition to shape hardware/software interfaces, programming models, and register/firmware contracts before tape-out.
  • Define the software strategy for multi-GPU and rack-scale topologies, including Infinity Fabric / UALink-style interconnect, collective communication (RCCL), memory coherence, and address translation across the platform.
  • Establish architecture for reliability, availability, and serviceability at scale, error detection, containment, telemetry, recovery, and graceful degradation across large clusters.
  • Set direction on performance: identify bottlenecks in the launch path, memory subsystem, and communication path, and define the software mechanisms to close them.
  • Produce architecture specifications, reference designs, and design reviews that align firmware, driver, runtime, and framework teams onto a shared plan.
  • Act as a technical anchor across AMD and with strategic hyperscale and AI customers - translating their workload requirements into architectural direction and representing AMD in deep technical engagements., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

Requirements

  • Linux Memory Management and Heterogeneous Memory Management (HMM).
  • GPU / DRM driver development.
  • Cache coherence and memory consistency protocols.
  • GPU Networking technologies including including scale up transport, NVLink, UALink, RDMA, and peer-direct.
  • Scale-up and scale-out networking, and communication collectives (e.g., RCCL/NCCL), MPI, or SHMEM.
  • Data-center fabrics such as Infinity Fabric, UALink, Ultra Ethernet, InfiniBand, and RoCE, with topology-aware software., * Direct experience with the ROCm stack, AMD Instinct, CUDA, or comparable GPU compute ecosystems.
  • Hands-on experience developing or optimizing GPU compute kernels (HIP, CUDA, Triton, or assembly-level tuning).
  • Experience with deep learning frameworks like PyTorch and TensorFlow in particular including framework integration, custom operators, and performance tuning on GPU backends.
  • Familiarity with the broader ML framework and compiler ecosystem (PyTorch, JAX, TensorFlow, ONNX, MLIR/compiler stacks).
  • Hands-on experience with GPU compute technologies such as OpenCL and Vulkan.
  • Familiarity with AI/ML training and inference workloads (transformers, large-scale distributed training, KV-cache and memory pressure, inference serving).
  • Background in virtualization, multi-tenancy, confidential computing, or cloud GPU provisioning.
  • Contributions to open-source kernel, driver, or runtime projects.

ACADEMIC CREDENTIALS:

  • Bachelor’s or Master’s in Electrical Engineer, Computer Engineering, Computer Science, or a closely related field

About the company

At AMD, our mission is to build great products that accelerate next-generation computing experiences-from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges-striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:42 min

Navigating emerging hardware standardization in vendor programming ecosystems

Paul Graham Paul Graham · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all