> Markdown version of [/jobs/ext/2720430-director-of-infrastructure-engineering](https://www.wearedevelopers.com/jobs/ext/2720430-director-of-infrastructure-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Director of Infrastructure Engineering - **Company:** RUNPOD INC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $225,000.0 - $325,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Border Gateway Protocol, Computer Clusters, DevOps, File Systems, Distributed Data Store, Distributed Systems, Ethernet, General-Purpose Computing on Graphics Processing Units, InfiniBand, Routing, Performance Tuning, Remote Direct Memory Access, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Database Engines, Weka, Ceph (Software), High Performance Computing, Deep Learning, Mttr, Kubernetes, Low Latency, Bare Metal, Free and Open-Source Software, Terraform, Nvme - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/director-of-infrastructure-engineering-runpod-8966712 ## About the Role * Engineering Leadership Experience: 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments. * Deep Infrastructure Expertise: 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms. * HPC & Advanced Networking: Proven hands-on background or strong architectural understanding of ultra-low latency networking. Deep familiarity with InfiniBand and/or RoCE, spine-leaf architectures, and global WAN routing protocols (BGP). * Storage Systems Knowledge: Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) capable of handling heavy AI/ML I/O loads. * SRE / DevOps Culture: Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks. * Remote-First Operating Excellence: Experience building culture, accountability, and momentum across distributed technical teams. * Communication & Collaboration: Clear written and verbal communication, strong stakeholder management, and calm, decisive leadership during high-stakes operational incidents. * Background Check: Successful completion of a background check., * Direct experience architecting and operating infrastructure specifically optimized for massive GPU clusters and AI/ML workloads. * Deep understanding of hardware architectures, GPU interconnects (NVLink), and datacenter topology. * Track record of scaling infrastructure teams in hyper-growth startup environments. * Open-source contributions or active recognition within the infrastructure, networking, or Kubernetes communities. ## Description We're looking for a Director of Infrastructure Engineering to lead and scale Runpod's core cloud and bare-metal environments. This role owns the critical foundational layers of our platform-Site Reliability Engineering (SRE), global networking, High-Performance Computing (HPC) networks, and distributed storage engines. You'll build the operating rhythm, culture, and technical direction that ensures Runpod remains highly available, performant, and capable of scaling to meet massive GPU computing demands. You will partner tightly with Product Engineering, Product, and GTM leadership to support enterprise customers. Your focus is ensuring that our underlying infrastructure provides the absolute fastest, most reliable, and lowest-latency path for customers training and running large-scale AI workloads. Responsibilities: * Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage. Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation. * Architect HPC & Global Networking: Oversee the design, scaling, and operation of Runpod's global network backbone, as well as ultra-low-latency HPC cluster networks. Drive the implementation and optimization of InfiniBand and RDMA over Converged Ethernet (RoCE) to support massive, multi-node GPU training workloads. * Drive Storage Engine Innovation: Direct the architecture and performance tuning of highly scalable, distributed storage systems. Ensure our storage engines can deliver the massive IOPS and throughput required to keep high-end GPUs fed with data during deep learning tasks. * Build a High-Output Org: Hire, mentor, and grow highly technical engineering managers and senior ICs (network architects, systems engineers, SREs). Create a culture of ownership, operational excellence, and craft in a remote-first environment. * Translate Scale into Strategy: Partner with Program Management and Product to forecast capacity requirements, shape technical roadmaps, and convert massive scale challenges into clear technical scopes, milestones, and measurable outcomes. * Continuously Improve Systems & Flow: Drive measurable improvements in infrastructure reliability and delivery metrics, such as deployment frequency, MTTR (Mean Time To Recovery), infrastructure as code (IaC) coverage, and system uptime. * Architectural Stewardship: Provide architectural oversight for bare-metal provisioning, virtualization layers, network fabrics, and storage clusters, ensuring seamless scalability without becoming a bottleneck for your teams. * Cross-Functional Partnership: Coordinate cleanly with product delivery and platform teams to ensure the infrastructure primitives they rely on are robust, well-documented, and highly available. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers)