AI Research Scientist - Infrastructure Engineer, Reinforcement Learning

Advanced Micro Devices, Inc.
Santa Clara, CA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$204,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence C++ (Programming Language) Software Debugging Python (Programming Language) Machine Learning Performance Tuning Software Systems Management of Software Versions Reinforcement Learning Pytorch Large Language Models Containerization
+1 more
Information Technology

Job description

  • Design and implement distributed RL training stacks (data parallel, pipeline parallel, or hybrid) integrated with AMD’s schedulers and storage
  • Build high-throughput rollout workers, trajectory stores, and reward computation pipelines with versioning and audit trails
  • Instrument jobs for debugging (NaNs, stragglers, OOMs), implement autoscaling and preemption-safe checkpointing
  • Collaborate with research scientists on experiment templates, hyperparameter sweeps, and safe promotion paths from research to wider team use
  • Drive reliability: on-call rotations, runbooks, and postmortems for infra incidents affecting RL training, AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

Requirements

You profile before you optimize; you treat researcher time as expensive as GPU time. You communicate SLAs, capacity plans, and incident patterns clearly and partner on cost-quality tradeoffs., * Strong systems track record in machine learning (ML) platforms with deep systems expertise and demonstrated technical impact.

  • Deep experience with PyTorch (or JAX), NCCL/MPI-style distributed training, and GPU cluster orchestration
  • Prior ownership of RL training infra, LLM post-training pipelines, or large-scale experiment management
  • Proficiency in C++/Python performance tuning, I/O optimization, and containerized workloads, * Bachelor’s degree required; Master’s or PhD preferred in Computer Science for research-heavy collaboration depth is preferred.

About the company

At AMD, our mission is to build great products that accelerate next-generation computing experiences-from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges-striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.localjobnetwork.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:54 min

Leveraging chip sets and software for artificial intelligence efficiency

Ankit Patel Ankit Patel ¡ WWC 2025

2:35 min

Preventing remote code execution in PyTorch models

Balåzs Kiss ¡ WWC 2023

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

2:36 min

Experiencing Conway's Law in heterogeneous software systems

Michael Jaeger Michael Jaeger ¡ Europe 2026 Virtual

2:03 min

Solving complex engineering challenges in artificial intelligence deployment

Nico Axtmann ¡ WWC 2022

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 ¡ WWC Europe 2026

Videos

See all

Related articles

See all