Fellow, Platform Power Management Architect - AMD Instinct GPUs

Advanced Micro Devices, Inc.
San Jose, CA, United States
3 months ago
Apply on careerarc.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Advanced Configuration and Power Interface (ACPI) Artificial Intelligence Computer Clusters Computer Engineering Data Centers Linux Network Interface Controllers Firmware Linux Kernel Graphics Processing Unit (GPU) High Performance Computing Data Center Networking
+1 more
Information Technology

Job description

WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next-generation computing experiences-from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges-striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.

The Role

AMD is seeking a Platform Power Management Architect to define and drive end-to-end power architecture for AMD Instinct data center GPU platforms. This role is responsible for system-level power strategy, spanning silicon capabilities, board-level power delivery, firmware, Linux power management, and rack-scale deployment considerations.

The architect will work cross-functionally with silicon, firmware, platform hardware, Linux kernel/ROCm software, and data center system teams to optimize performance per watt, ensure power integrity and reliability, and deliver accurate power projections for current and next-generation Instinct platforms. Exposure to scale-up or scale-out networking fabrics is highly desirable.

Key Responsibilities

Platform Power Architecture

  • Define the end-to-end power management architecture for AMD Instinct data center GPUs, spanning silicon, package, board, system, and rack levels.
  • Own platform-level power concepts including power states, power limits, throttling policies, telemetry, and power-performance trade-offs.
  • Act as the technical authority for power-related architectural decisions across multiple Instinct programs.

Power Delivery & Rail Optimization

  • Lead power rail architecture and optimization, including rail partitioning, sequencing, voltage/frequency domains, and efficiency trade-offs.
  • Partner with hardware and silicon teams to optimize VR efficiency, transient response, and steady-state power delivery under AI/HPC workloads.
  • Influence silicon and platform features to improve power scalability and robustness across SKUs and deployment configurations.

Linux & Software Power Management

  • Define requirements and architecture for Linux-based power management, including interactions with kernel frameworks, drivers, firmware, and ROCm components.
  • Collaborate with software teams on power telemetry, control interfaces, policy enforcement, and observability.
  • Ensure alignment between platform power capabilities and software-visible controls for data center operators.
  • Solid understanding of RTPM, ACPI and Suspend to Idle / S0ix flows.

Power Modeling & Projections

  • Develop and own power projection methodologies for GPUs, platforms, and multi-GPU systems across representative workloads.
  • Provide power projections and sensitivity analyses to support product planning, system design, customer engagements, and thermal/rack planning.
  • Validate projections against lab data and silicon characterization results, closing gaps between model and reality.

System & Fabric Awareness

  • Incorporate scale-up (e.g., high-bandwidth GPU interconnects) and scale-out (e.g., networking fabrics) considerations into platform power strategy.
  • Understand and influence the power impact of interconnects, NICs, switches, and fabric topologies in large GPU clusters.
  • Partner with fabric and system architects to ensure coherent power budgeting at node and rack scale.

Cross-Functional Leadership

  • Drive alignment across silicon, firmware, hardware, Linux, ROCm, platform, and data center solution teams.
  • Produce clear architectural documentation, power models, and executive-level summaries.
  • Represent platform power architecture in technical reviews with senior leadership and external partners.

Required Qualifications

  • Expert-level background in platform, system, or silicon architecture with significant focus on power management.
  • Strong understanding of power delivery networks (PDN), voltage regulation, rail optimization, and power integrity fundamentals.
  • Hands-on experience with Linux power management concepts, kernel/driver interactions, or system-level power control.
  • Experience building or consuming power models and projections for complex systems.
  • Ability to work across hardware and software boundaries and influence architectural decisions.
  • Bachelor’s degree in Electrical Engineering, Computer Engineering, Computer Science, or related field (Master’s or PhD preferred).

Preferred Qualifications

  • Experience with data center GPUs, accelerators, or high-performance SoCs.
  • Exposure to scale-up GPU fabrics and/or scale-out data center networking.
  • Familiarity with telemetry, power capping, workload-aware power management, or fleet-level power optimization.
  • Experience presenting architectural trade-offs to senior technical leadership.
  • Background in HPC or AI training/inference systems.

This role is not eligible for visa sponsorship.

LI-BW2

LI-HYBRID

This role is not eligible for visa sponsorship.

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Requirements

  • Expert-level background in platform, system, or silicon architecture with significant focus on power management.
  • Strong understanding of power delivery networks (PDN), voltage regulation, rail optimization, and power integrity fundamentals.
  • Hands-on experience with Linux power management concepts, kernel/driver interactions, or system-level power control.
  • Experience building or consuming power models and projections for complex systems.
  • Ability to work across hardware and software boundaries and influence architectural decisions.
  • Bachelor’s degree in Electrical Engineering, Computer Engineering, Computer Science, or related field (Master’s or PhD preferred)., * Experience with data center GPUs, accelerators, or high-performance SoCs.
  • Exposure to scale-up GPU fabrics and/or scale-out data center networking.
  • Familiarity with telemetry, power capping, workload-aware power management, or fleet-level power optimization.
  • Experience presenting architectural trade-offs to senior technical leadership.
  • Background in HPC or AI training/inference systems.

About the company

At AMD, our mission is to build great products that accelerate next-generation computing experiences-from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges-striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on careerarc.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:18 min

Replacing legacy CPUs to reduce data center energy consumption

Thomas Schmidt Thomas Schmidt · World Congress 2024

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · World Congress 2024

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all