Sr. Data Center Systems Engineer
Role details
Job location
Tech stack
Job description
Join AMD's Data Center Platform Engineering Group (DPEG) and play a critical role in enabling the reliability, deployment, and availability of next-generation data center platforms. In this highly collaborative, hands-on engineering role, you will work across hardware, firmware, software, and systems teams to diagnose complex technical challenges and ensure platform stability at scale. You will have the opportunity to drive root-cause investigations, validate cutting-edge server technologies, and partner with internal engineering teams, OEMs, ODMs, and vendors to resolve critical issues. This role offers exposure to advanced CPU, GPU, and platform technologies while providing the opportunity to influence product quality, system performance, and operational excellence. THE PERSON: The ideal candidate is a highly motivated and self-driven engineer with a passion for solving complex technical problems. They thrive in a fast-paced environment, possess strong debugging and analytical skills, and can effectively collaborate across multiple teams and disciplines. Success in this role requires excellent communication skills, strong ownership, attention to detail, and the ability to drive issues to resolution while balancing multiple priorities., * Drive day-to-day platform support activities, including system availability, workload failures, thermal, power, networking, clustering, and capacity-related issues.
- Develop and execute debug plans, methodologies, and strategies to identify root causes and resolve complex system-level issues.
- Coordinate and track issue resolution efforts across engineering teams while ensuring timely communication and escalation.
- Collaborate with hardware, firmware, software, reliability, and validation teams to define requirements and assess technical risks.
- Develop and execute laboratory validation plans to ensure platform functionality, reliability, and performance.
- Partner with OEMs, ODMs, vendors, and internal stakeholders to investigate issues and drive corrective actions.
- Apply system-level reliability knowledge to improve platform quality and operational efficiency.
- Drive continuous improvement initiatives related to debug processes, validation methodologies, automation, and overall engineering effectiveness.
- Contribute to the development of scalable support and validation processes that improve long-term platform reliability and availability., AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. This posting is for an existing vacancy.
Requirements
- Experience debugging complex hardware, firmware, or system-level issues and driving root-cause investigations.
- Knowledge of server, CPU, GPU, SoC, microcontroller, or data center platform architectures.
- Experience with hardware validation, verification, and reliability methodologies.
- Familiarity with Linux environments and system debug techniques.
- Knowledge of BMC, BIOS, FPGA, and CPLD technologies.
- Experience using issue tracking and project management tools such as JIRA.
- Scripting and test automation experience.
- Understanding of high-speed industry-standard interfaces and protocols, including PCIe.
- Experience collaborating with OEMs, ODMs, suppliers, and external partners.
- Familiarity with laboratory equipment and hardware bring-up environments.
- Strong analytical, troubleshooting, and problem-solving abilities.
- Excellent verbal and written communication skills., * Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or a related technical field preferred.
- Master's degree in a relevant engineering discipline desired.
- Relevant industry certifications or specialized training in systems engineering, validation, or reliability engineering are beneficial.