> Markdown version of [/jobs/ext/458147-infrastructure-services-operations-lead](https://www.wearedevelopers.com/jobs/ext/458147-infrastructure-services-operations-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Services Operations Lead - **Company:** Advanced Micro Devices, Inc. - **Location:** Austin, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, BIOS, C++ (Programming Language), Cloud Engineering, Computer Clusters, Configuration Management, Computer Engineering, Data Centers, Distributed Data Store, Distributed Systems, Infrastructure as a Service (IaaS), InfiniBand, Python (Programming Language), Machine Learning, Ansible, Software Engineering, Weka, Ceph (Software), High Performance Computing, Technical Debt, Infrastructure Automation Frameworks, Storage Technologies, Information Technology, SDN Network, Bare Metal, Build Tools, Slurm, Terraform - **Published:** June 7, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=45358c3aa07fac13 ## About the Role * Architectural Leadership: Proven track record as a Principal Engineer or Architect designing large-scale (10,000+ node) distributed systems or private/public cloud environments. * Bare-Metal & IaaS: Deep expertise in building IaaS layers, including bare-metal provisioning (e.g., Ironic, Tinkerbell, or custom PXE/Redfish workflows) and Software-Defined Networking (SDN). * GPU/HPC Expertise: Technical understanding of high-density GPU platforms (AMD Instinct or similar), InfiniBand/RoCE fabrics, and the specific architectural requirements of AI/ML training at scale. * Storage Architecture: Hands-on experience architecting high-performance distributed storage solutions (e.g., Weka, Lustre, Ceph) for massive datasets. * Cloud-Native & Orchestration: Expert-level knowledge of Kubernetes internals, SLURM, and how to bridge traditional HPC scheduling with modern container orchestration. * Tooling & Languages: Proficiency in Go, Python, or C++, and deep experience with Terraform, Ansible, or custom-built controllers/operators. * Global Consolidation: Experience leading the technical migration from decentralized, heterogeneous data center silos to a unified, global hardware/software standard. ACADEMIC CREDENTIALS: Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related technical field preferred Location could be in: Austin, TX, Santa Clara, CA, Seattle, WA, or Markham, Canada ## Description We are seeking a Principal Engineer to serve as the Lead Architect for AMD's internal infrastructure platform supporting Instinct GPU development and deployment. In this role, you will be the primary technical authority responsible for the "how" behind AMD's global data center transformation. You will design the integrations that transition our infrastructure from manual, high-touch management to a highly automated, software-defined environment. You will operate at the highest technical level, bridging the gap between hardware capabilities and software-defined control planes. Your mission is to eliminate technical debt by architecting global, consolidated processes that allow our infrastructure to scale without a linear increase in headcount. This role culminates in the design and delivery of an internal Infrastructure-as-a-Service (IaaS) offering-providing standardized compute, storage, and networking APIs across our global fleet. THE PERSON: You are a visionary technologist who views infrastructure through the lens of software engineering. You don't just manage systems; you build systems that manage systems. You have a deep understanding of the full stack-from silicon and high-density power requirements to kernel-level networking and distributed storage protocols. You are a "systems thinker" who can identify fragmented global processes and consolidate them into elegant, automated workflows. You thrive on solving the structural "how" of complex problems, such as bare-metal provisioning at scale and global state consistency. You communicate with technical authority, mentoring senior engineers while providing executive leadership with a clear technical roadmap for a consolidated, scalable future., * Architect the IaaS Foundation: Design the technical specifications and architecture for a unified internal IaaS offering (Compute, Storage, Networking) to be deployed across global datacenters within the first year. * Solve for "How": Define the technical standards and implementation patterns for automated provisioning, configuration management, and hardware lifecycle orchestration. * Reduce Tech Debt: Conduct deep-dive audits of existing fragmented workflows; design and implement consolidated, global technical standards to replace legacy, site-specific manual processes. * Enable Human-Resource Scalability: Engineer automation frameworks that decouple fleet growth from headcount growth, ensuring a small team can manage massive, heterogeneous GPU clusters. * Define the Technical Stack: Lead the selection and architectural integration of technologies (e.g., Bare Metal-as-a-Service, SDN, Distributed Storage) that form the backbone of the Instinct development platform. * Drive Infrastructure-as-Code (IaC): Establish the engineering patterns for the entire infrastructure lifecycle, ensuring every component from BIOS settings to network fabric is version-controlled and programmatically accessible. * Cross-Functional Technical Leadership: Act as the final technical arbiter between Platform Engineering, SRE, and Ops to ensure architectural alignment and the removal of technical bottlenecks. * Performance Engineering: Optimize the interaction between Instinct GPU hardware and the underlying infrastructure to ensure maximum utilization and reliability for AI/HPC workloads. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [How to Submit CFPs and Get into Public Speaking - Moran Weber](https://www.wearedevelopers.com/videos/2112-how-to-submit-cfps-and-get-into-public-speaking-moran-weber) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Top Must-Visit Developer Conferences in the US in 2026](https://www.wearedevelopers.com/magazine/679-top-must-visit-developer-conferences-in-the-us-in-2026) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)