> Markdown version of [/jobs/ext/1803380-sr-director-it](https://www.wearedevelopers.com/jobs/ext/1803380-sr-director-it). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr Director IT - **Company:** Advanced Micro Devices, Inc. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $231,280.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Cloud Computing, Computer Clusters, Complex Networks, Computer Engineering, Data Centers, Extract Transform Load (ETL), Linux, Ethernet, Network Interface Controllers, Routing, Performance Tuning, Remote Direct Memory Access, Cloud Services, Systems Integration, Private Cloud Environment, High Performance Computing, Software Troubleshooting, HybridCloud, Backend, Information Technology, Front End Software Development, Hardware Infrastructure - **Published:** July 3, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/87623360/1 ## About the Role The position requires a hands-on technical leader with strong experience in AI/HPC infrastructure, high-speed networking, data center enablement, security, automation, and production operations. This leader will lead team and help scale the platform, operating model, and roadmap to support AMD's growing internal and external AI compute needs., The ideal candidate is a senior leader with deep technical expertise and strong execution skills. They have experience building and operating GPU clusters, enabling AI workloads, managing complex networks, improving utilization, and delivering reliable compute services. They can operate strategically while remaining hands-on when needed, and they bring a strong vision for how AI compute platforms should evolve., * 15+ years of experience in infrastructure engineering, cloud infrastructure, AI/HPC systems, networking, data center infrastructure, or related fields. * Proven experience building, scaling, and operating GPU clusters, AI compute platforms, HPC environments, or large-scale cloud infrastructure. * Strong understanding of GPU infrastructure, cluster operations, Linux, automation, orchestration, monitoring, and production support. * Deep experience with high-speed networking, including Ethernet fabrics, RoCE/RDMA, NICs, switches, optics, routing, segmentation, and performance troubleshooting. * Experience enabling storage and data movement for AI/HPC workloads. * Experience with private cloud, GPU-as-a-Service, hybrid cloud, or internal compute platform delivery. * Strong background in security, multi-tenancy, access control, and operational reliability. * Demonstrated ability to lead technical teams and communicate effectively with senior stakeholders. PREFERRED EXPERIENCE * Experience with AMD GPU platforms, ROCm, or accelerator software ecosystems. * Experience with direct liquid cooling, liquid-to-chip cooling, high-density AI racks, and data center readiness for GPU workloads. * Experience selecting or enabling data centers for AI/HPC infrastructure. * Experience integrating on-premises AI compute with neo-cloud and tier-1 cloud providers. * End-to-end understanding of AI infrastructure from data center and networking through workload execution and token delivery. ACADEMIC CREDENTIALS Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related technical field, or equivalent experience. Advanced degree preferred. ## Description AMD is seeking a Senior Director, AI Compute Infrastructure & GPU Cloud Services to lead the strategy, architecture, deployment, and operations of large-scale AI compute environments. This role will enable GPU clusters, private AI compute cloud platforms, GPU-as-a-Service capabilities, and hybrid deployments across on-premises infrastructure, neo-clouds, and tier-1 public clouds., * Define and execute the roadmap for AMD's AI compute infrastructure, including on-prem GPU clusters, private AI cloud, neo-cloud, and public cloud environments. * Lead GPU cluster bring-up, production operations, monitoring, automation, capacity planning, availability, and utilization improvement. * Build and operate GPU-as-a-Service capabilities for internal users and strategic external needs. * Architect and manage front-end and back-end networks for AI workloads, including high-speed egress, storage connectivity, RoCE/RDMA fabrics, congestion management, and performance tuning. * Partner with storage, security, data center, platform, software, and cloud teams to deliver secure, scalable, and reliable AI compute services. * Drive end-to-end infrastructure readiness from data center planning through workload execution and token delivery. * Support data center evaluation and enablement for high-performance AI compute, including power, cooling, rack density, and direct liquid cooling requirements. * Lead, mentor, and grow a small technical team while working cross-functionally with engineering, IT, vendors, and executive stakeholders., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)