Principal System Debug Engineer

Advanced Micro Devices, Inc.
San Jose, CA, United States
10 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

C (Programming Language) Microsoft Excel Artificial Intelligence Computing Platforms Automation of Tests Intelligent Platform Management Interface BIOS C++ (Programming Language) Computer Programming Computer Engineering Microarchitecture Software Debugging
+23 more
Linux Device Drivers Microprocessors Ethernet Firmware Hardware Design Hyper-V Hypervisor Joint Test Action (IEEE Standards) Python (Programming Language) Kernel-Based Virtual Machine PCI Express Cloud Services Signal Integrity Software Engineering System Software WinDBg Extensible Firmware Interface Information Technology Windows Kernel Pcb Layout Hardware Debugging Vmware

Job description

As a Principal System Debug Engineer, you will lead the system-level debug of our advanced microprocessors within complex customer applications and system design environments. You will act as the critical bridge between our internal silicon/software engineering teams and our top-tier OEM, ODM, and Hyperscaler customers., * Technical Leadership & Debug: Lead the triage, investigation, and root-cause analysis of complex system-level issues involving CPU silicon, platform hardware, firmware, drivers and OS/hypervisor interactions in customer environments.

  • System Bring-Up: Drive early platform bring-up activities for next-generation CPUs, both in internal labs and on-site at customer facilities, ensuring rapid time-to-market.
  • Customer Engagement: Serve as the primary technical liaison for strategic customers. Guide customer engineering teams through platform design, debug methodologies, and issue resolution.
  • Cross-Functional Collaboration: Partner closely with internal Silicon Design, Architecture, Firmware (BIOS/UEFI/BMC), OS/Kernel, and Validation teams to drive systemic fixes and influence future CPU architectures based on customer feedback.
  • Debug Methodology & Tooling: Architect and develop advanced debug methodologies, scripts, and tools to accelerate issue isolation across the hardware/software boundary.
  • Escalation Management: Act as the technical task force leader during critical customer escalations, providing clear executive updates and driving technical action plans under tight deadlines., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

Requirements

You are a subject matter expert and strong technical contributor with strong processor architecture, hardware and software expertise, and extensive system level debug experience. You excel as part of a team where technical leadership, customer engagement, communication and team skills are highly valued., * CPU Architecture: Deep, foundational knowledge of modern CPU architectures (x86 mandatory), including pipelines, cache coherency, memory controllers, and power management (C-states, P-states). RAS features, MCA (Machine Check Architecture), SoC and platform security features, RoT (root-of-trust).

  • I/O & Interconnects: Extensive expertise in the architecture and protocol-level debug of high-speed I/O interfaces, including PCIe (Gen 4/5/6), CXL, DDR4/DDR5, Ethernet, USB, and low-speed buses (I2C, SPI, I3C, eSPI).
  • Firmware & Software Stacks: Strong understanding of the full system software stack, including BIOS/UEFI, BMC/IPMI, Linux/Windows kernel internals, device drivers, and hypervisors (KVM, VMware, Hyper-V).
  • Platform Design: Solid grasp of system-level hardware design, including board schematics, PCB layout, power delivery networks (VRMs), clocking, and signal integrity fundamentals.
  • Hands-On Debug: Proven track record of performing complex system bring-up and debug using hardware tools (JTAG/ITP debuggers, oscilloscopes, logic analyzers, protocol analyzers) and software debuggers (GDB, WinDbg, kernel panics/crash dump analysis).
  • Experience working directly with Cloud Service Providers (Hyperscalers) and Tier-1 Server/Client OEMs.
  • Proficiency in scripting and programming languages (Python, C, C++) for test automation and debug tool development.
  • Hands-on expertise in debugging hardware (system bring-up, signal integrity, power integrity issues) and debugging firmware/software running on actual hardware platforms.
  • Experience with RAS (Reliability, Availability, and Serviceability) architecture and machine check exception (MCE) analysis.

ACADEMIC CREDENTIALS:

Bachelor’s, Master’s, or PhD in Electrical Engineering, Computer Engineering, Computer Science, or a related field.

About the company

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger - technology that moves the world forward. Join us and, together, we’ll advance your career.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:20 min

Certifying safety-critical automotive software and deploying progressive testing strategies

David Romić · World Congress 2023

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 · LIVE

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · World Congress 2024

Videos

See all

Related articles

See all