Sr Director IT
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
AMD is seeking a Senior Director, AI Compute Infrastructure & GPU Cloud Services to lead the strategy, architecture, deployment, and operations of large-scale AI compute environments. This role will enable GPU clusters, private AI compute cloud platforms, GPU-as-a-Service capabilities, and hybrid deployments across on-premises infrastructure, neo-clouds, and tier-1 public clouds., * Define and execute the roadmap for AMDâs AI compute infrastructure, including on-prem GPU clusters, private AI cloud, neo-cloud, and public cloud environments.
- Lead GPU cluster bring-up, production operations, monitoring, automation, capacity planning, availability, and utilization improvement.
- Build and operate GPU-as-a-Service capabilities for internal users and strategic external needs.
- Architect and manage front-end and back-end networks for AI workloads, including high-speed egress, storage connectivity, RoCE/RDMA fabrics, congestion management, and performance tuning.
- Partner with storage, security, data center, platform, software, and cloud teams to deliver secure, scalable, and reliable AI compute services.
- Drive end-to-end infrastructure readiness from data center planning through workload execution and token delivery.
- Support data center evaluation and enablement for high-performance AI compute, including power, cooling, rack density, and direct liquid cooling requirements.
- Lead, mentor, and grow a small technical team while working cross-functionally with engineering, IT, vendors, and executive stakeholders., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMDâs âResponsible AI Policyâ is available here.
Requirements
The position requires a hands-on technical leader with strong experience in AI/HPC infrastructure, high-speed networking, data center enablement, security, automation, and production operations. This leader will lead team and help scale the platform, operating model, and roadmap to support AMDâs growing internal and external AI compute needs., The ideal candidate is a senior leader with deep technical expertise and strong execution skills. They have experience building and operating GPU clusters, enabling AI workloads, managing complex networks, improving utilization, and delivering reliable compute services. They can operate strategically while remaining hands-on when needed, and they bring a strong vision for how AI compute platforms should evolve., * 15+ years of experience in infrastructure engineering, cloud infrastructure, AI/HPC systems, networking, data center infrastructure, or related fields.
- Proven experience building, scaling, and operating GPU clusters, AI compute platforms, HPC environments, or large-scale cloud infrastructure.
- Strong understanding of GPU infrastructure, cluster operations, Linux, automation, orchestration, monitoring, and production support.
- Deep experience with high-speed networking, including Ethernet fabrics, RoCE/RDMA, NICs, switches, optics, routing, segmentation, and performance troubleshooting.
- Experience enabling storage and data movement for AI/HPC workloads.
- Experience with private cloud, GPU-as-a-Service, hybrid cloud, or internal compute platform delivery.
- Strong background in security, multi-tenancy, access control, and operational reliability.
- Demonstrated ability to lead technical teams and communicate effectively with senior stakeholders.
PREFERRED EXPERIENCE
- Experience with AMD GPU platforms, ROCm, or accelerator software ecosystems.
- Experience with direct liquid cooling, liquid-to-chip cooling, high-density AI racks, and data center readiness for GPU workloads.
- Experience selecting or enabling data centers for AI/HPC infrastructure.
- Experience integrating on-premises AI compute with neo-cloud and tier-1 cloud providers.
- End-to-end understanding of AI infrastructure from data center and networking through workload execution and token delivery.
ACADEMIC CREDENTIALS
Bachelorâs degree in Computer Science, Electrical Engineering, Computer Engineering, or a related technical field, or equivalent experience. Advanced degree preferred.
About the company
At AMD, our mission is to build great products that accelerate next-generation computing experiences-from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, youâll discover the real differentiator is our culture. We push the limits of innovation to solve the worldâs most important challenges-striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on jobs.localjobnetwork.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
7 Cloud Computing Trends Coming in 2025 for Developers
What Industries Outside of AI Are Hiring The Most AI Experts?
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud