> Markdown version of [/jobs/ext/2805415-dcgpu-platform-system-manager](https://www.wearedevelopers.com/jobs/ext/2805415-dcgpu-platform-system-manager). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DCGPU Platform System Manager - **Company:** Advanced Micro Devices, Inc. - **Location:** Austin, TX, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cluster Analysis, Data Centers, Software Debugging, System Availability - **Published:** September 9, 2026 - **Apply:** https://www.careerarc.com/job-listing/amd-jobs-dcgpu-platform-system-manager-52866153 ## About the Role Experienced, self-motivated individual that has previously managed a data center preferably in the GPU system space w/500+ systems/platform. Person should be able to communicate updates on the state of the fleet, ensure the team is working on deployment/availability of the fleet, as well be able to solve technical issues that arise., * Experience managing GPU data center employees, systems, day-to-day activities * Knowledge in solving issues around capacity planning, power, thermal, networking, clustering * Executive-level focused communication * Management experience in GPU data centers that have a vast array of different systems/platforms * Co-work with external stakeholders/vendors and ability to openly communicate/drive issues * Collaborate with internal stakeholders on root causing issues, driving issue meetings * Hands-on experience with day-to-day issues that arise in a data center (capacity, network, power, thermal) * Leadership and communication skills that require presentations in executive forum * Ability to clearly articulate the work by the team (ins/outs), needs, and recruit necessary skills ACADEMIC CREDENTIALS: * Bachelors degree in engineering ## Description WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences-from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you'll discover the real differentiator is our culture. We push the limits of innovation to solve the world's most important challenges-striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. DCGPU PLATFORM SYSTEM MANAGER THE ROLE: The Data Center Platform Engineering Group (DPEG) Manager that will be the primary on-site leader for a data center in north Austin, Texas. This individual will need to be on-site daily driving day-to-day activities that will include the deployment and availability of highly complex AMD Instinct platforms that will be the backbone of the work required to release state-of-the-art technologies. This manager will have personnel responsibilities and be required to manage the work required to run a data center with well over 1000 systems. THE PERSON: Experienced, self-motivated individual that has previously managed a data center preferably in the GPU system space w/500+ systems/platform. Person should be able to communicate updates on the state of the fleet, ensure the team is working on deployment/availability of the fleet, as well be able to solve technical issues that arise. KEY RESPONSIBILITIES: * On-site management of a data center site with a vast array of differing platforms/system * Personnel and work assigned management of data center engineers and technicians * Clear / Concise communication in open daily meetings on "high attention" given to system availability * Knowledge in requirements for root causing issues and understanding/mapping debug methodologies * Provide leadership input/recommendations for improvements / help drive organizational initiatives PREFERRED EXPERIENCE: * Experience managing GPU data center employees, systems, day-to-day activities * Knowledge in solving issues around capacity planning, power, thermal, networking, clustering * Executive-level focused communication * Management experience in GPU data centers that have a vast array of different systems/platforms * Co-work with external stakeholders/vendors and ability to openly communicate/drive issues * Collaborate with internal stakeholders on root causing issues, driving issue meetings * Hands-on experience with day-to-day issues that arise in a data center (capacity, network, power, thermal) * Leadership and communication skills that require presentations in executive forum * Ability to clearly articulate the work by the team (ins/outs), needs, and recruit necessary skills ACADEMIC CREDENTIALS: * Bachelors degree in engineering This role is not eligible for visa sponsorship. #LI-TL1 Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. This posting is for an existing vacancy. ## Related Videos - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Using AI Without Losing Your Skills](https://www.wearedevelopers.com/videos/2045-using-ai-without-losing-your-skills) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters](https://www.wearedevelopers.com/magazine/571-dev-digest-162-ai-careers-mcp-aws-best-practices-floppy-sweaters) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)