> Markdown version of [/jobs/ext/1708451-director-engineering-global-ml-scheduling-infrastructure](https://www.wearedevelopers.com/jobs/ext/1708451-director-engineering-global-ml-scheduling-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Director, Engineering, Global ML Scheduling Infrastructure - **Company:** Google LLC - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Salary:** $307,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Data Centers, Job Scheduling, Machine Learning, Software Engineering, Graphics Processing Unit (GPU), Google Cloud, Cloud Platform System, Information Technology, Microservices - **Published:** July 11, 2026 - **Apply:** https://dejobs.org/x/x/88E4EA7703E14003B3B698C89B5FB193/job/ ## About the Role * Bachelor's degree in Computer Science or equivalent practical experience. * 15 years of experience in software engineering. * 10 years of experience managing and leading large-scale distributed engineering teams. * Experience managing international teams and driving cross-site organizational alignment. * Experience leading infrastructure engineering organizations, specifically managing control planes, cluster management systems, or distributed job scheduling platforms., * Experience leading enterprise-level AI transformations, including scaling TPU/GPU accelerator infrastructure, and accelerating the transition of complex ML research innovations into high-performance, production-ready developer platforms. * Ability to navigate complex matrixed organizations, influence technical strategy at the industry level, and drive convergence across legacy and modernized stacks. * Domain expertise in distributed resource management, machine learning training infrastructure, hardware accelerator orchestration (GPUs/TPUs), and large-scale cloud computing platforms. * Technical expertise in designing, building, and operating global-scale scheduling and orchestration systems, specifically specializing in multi-cell/multi-tenant scheduling ecosystems, throughput-oriented batch workloads, and resource optimization. ## Description A core suite of systems and infrastructure manages Google's global orchestration for throughput-oriented workloads across various fleet locations, maximizing resource efficiency on a massive scale. Specializing in accelerator scheduling and massive-scale Machine Learning (ML) training, this infrastructure serves both internal Google fleets and the Google Cloud Platform (GCP). Its capabilities are continuously expanding to encompass the entire traditional compute fleet alongside the accelerator fleet. By offering workload flexibility across spatial, platform, and quota dimensions, these systems achieve exceptionally high fleet occupancy while maintaining robust usability and reliability for third-party customers and all major Google product areas. As the Director of Engineering for the Global ML Scheduling Infrastructure team, you will lead the strategic direction, engineering execution, and operational excellence of Google's global orchestration layer. Leading a distributed organization of approximately 90 engineers across the US and Poland, you will oversee the mission-critical multi-cell scheduling ecosystem that powers Google's large-scale ML training, inference, and general throughput-oriented batch workloads. Collaborating with platform, storage, data center, networking, and resource management teams, you will drive new capabilities and support the growth and efficient usage of Google's fleet. You will partner with leads from Google product areas, such as Deepmind, Search, Ads, and YouTube, to accelerate the transition of research innovations to production, with focus on developer experience and acceleration of experimentation and productionization time. You will also contribute to delivering GPUs and Google's advanced internal technology, TPUs, to external customers via Google's Cloud Compute Platform. You will advocate for architectural innovation, drive efficiency initiatives that directly impact Google's infrastructure footprint, and foster a high-performance culture across international sites. Google Cloud accelerates every organization's ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google's cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems. Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $307000 - $428000 (USD) + 30% bonus target + equity + benefits, * Define and execute the long-term technical vision and engineering strategy for the Global ML Scheduling Infrastructure functional area, ensuring highly scalable and efficient workload scheduling across Google's global fleet. * Manage and grow a high-performing engineering organization distributed across the United States and Poland, and collaborating with functions including SRE, PMO, and analytics teams. * Partner strategically with executive leadership and cross-functional stakeholders across technical infrastructure and product areas to align platform capabilities with Alphabet's accelerating demands for machine learning and throughput-oriented computing. * Drive architectural evolution and operational excellence across the suite of scheduling microservices, maintaining rigorous SLOs for queuing, fair sharing, cell-level actuation, and multi-tenant resource optimization. * Advocate for a collaborative and psychologically safe environment that prioritizes talent development, imports and exports top engineering talent, and exemplifies Google's core leadership principles. ## Related Videos - [The Cloud is Calling: Answer with In-Demand Skills](https://www.wearedevelopers.com/videos/945-the-cloud-is-calling-answer-with-in-demand-skills) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [Microservices: how to get started with Spring Boot and Kubernetes](https://www.wearedevelopers.com/videos/242-microservices-how-to-get-started-with-spring-boot-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)