Lead Systems Management Architect

Advanced Micro Devices, Inc.
Austin, TX, United States
15 days ago
Apply on jobs.localjobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$163,200.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Computing Platforms Computer Engineering Data Center Infrastructure Management (CIM) Software Debugging Embedded Software Firmware Open Source Technology Software Requirements Analysis Management of Software Versions High Performance Computing Information Technology
+2 more
U-Boot Server Operating Systems & Platforms

Job description

AMD is seeking a Lead Systems Management Architect to lead key aspects of the architecture and feature definition for the management of AMD Instinct accelerators and associated server platforms. This is a highly visible technical leadership role contributing to the overall rack-scale management architecture. You will collaborate with fellow architects, internal engineering teams, AMD partners, and end customers to develop novel features and create solutions that support future AMD products and integration with customer datacenter infrastructure. A deep understanding of the entire out-of-band (OOB) ecosystem, from the DCIM SW layer down to the managed components, will be critical for success in this role., * Lead key aspects of architecture and feature definition for the OOB management of AMD Instinct GPU accelerators and associated server platforms, encompassing health monitoring, power and thermal management, firmware lifecycle, inventory, and error handling, while ensuring these capabilities integrate coherently into the broader rack-scale management architecture.

  • Define and contribute to the standards-based interface strategy for GPU and server platform manageability using DMTF Redfish, PLDM, MCTP, and related specifications, balancing standards compliance with AMD-specific and OEM extension requirements.
  • Work with BMC and embedded firmware teams to define OOB management feature requirements, including Redfish schemas, sensor and inventory representations, eventing, firmware update flows, and debug workflows specific to GPU and server platform components.
  • Contribute to firmware management architecture for AMD Instinct accelerators and server platforms, covering in-band and out-of-band update flows, versioning, dependency management, activation strategies, and recovery mechanisms.
  • Engage directly with end customers, AMD partners, and ODM/OEMs to understand datacenter integration requirements, DCIM and orchestration software expectations, and operational workflows, translating these into concrete feature and interface requirements and guiding partners through implementation.
  • Partner with GPU/SoC architects, board and system architects, firmware and software teams, security/RAS, and validation to translate architecture into production-ready deliverables, and contribute to conformance and validation strategy for platform manageability.
  • Help shape future AMD Instinct platform roadmaps through customer engagement and field learnings, participate in relevant standards and open-source communities including DMTF and OpenBMC, and mentor engineers and architects across the organization., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

Requirements

The ideal candidate brings deep, hands-on expertise across the full OOB management stack - comfortable reasoning about DCIM and orchestration software requirements at one end and getting into the details of BMC firmware, platform interfaces, and device-level protocols at the other. You have a strong background in GPU/accelerator or server platform management and can translate that experience into clear architectural direction that works in practice, not just on paper. You are as effective writing a detailed interface specification or debugging a bring-up issue as you are presenting a roadmap proposal to senior leadership or walking a customer through a platform integration. You work well in environments where requirements are still taking shape, build credibility through technical depth rather than title, and make the teams around you better., * Expert-level experience in platform management architecture, server manageability, BMC or embedded firmware, or GPU/accelerator platform design, including significant time in architect or technical leadership roles delivering solutions for datacenter, cloud, AI, or HPC environments.

  • Proven understanding of the full OOB management stack, from DCIM platforms and datacenter orchestration frameworks through BMC firmware down to device-level management protocols, with the ability to reason clearly across every layer.
  • Deep knowledge of DMTF Redfish including schema design, OEM extension strategy, eventing, and update service; strong understanding of PLDM and MCTP for platform inventory, monitoring, control, and firmware update workflows; hands-on experience with OpenBMC architecture and services is strongly preferred.
  • Experience with firmware security concepts including secure boot, root of trust, firmware signing, attestation, and SPDM, combined with a track record of producing architecture specifications, product requirements, conformance plans, and validation strategies that drive execution across internal teams and external partners.
  • Experience engaging directly with customers or ODM/OEM partners to gather requirements, present architecture proposals, and drive alignment on platform management capabilities; familiarity with AMD server or GPU platforms, AI/HPC system design, or OCP-aligned rack architectures is a plus.

ACADEMIC CREDENTIALS:

Bachelor’s or Master’s degree (preferred) in Computer Science, Computer Engineering, Electrical Engineering, or equivalent

About the company

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.

AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger - technology that moves the world forward.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.localjobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:06 min

Moving from basic embedded software to system functionalities

Réka Leisztner Réka Leisztner · World Congress 2025

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

3:51 min

Technical skills and collaborative mindsets for mobility engineering roles

Georg Kühberger +1 · LIVE

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

1:14 min

Addressing automotive mission-critical safety in embedded software development

David Romić · World Congress 2023

Videos

See all

Related articles

See all