Director of Data Center Facilities - AI Infrastructure
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
We’re looking for a Director of Data Center Facilities - AI Infrastructure to join our team during an exciting phase of growth. In this role, you will own the physical infrastructure strategy, operations, reliability, and expansion of a high-density AI data center built to support GPU computing at scale.
You’ll make sure the facility can deliver the power, cooling, connectivity, and operational capacity that fast-growing AI compute platforms demand, while leading multidisciplinary facilities teams and partnering closely with Operations, Engineering, Network, Construction, Security, Finance, Procurement, utilities, and technology vendors.
This is a highly visible leadership role for someone with deep experience running mission-critical, high-density environments and a strong grasp of what AI workloads require, including extreme rack densities, liquid cooling, power quality, thermal management, and rapid deployment.and impact.
What You’ll Do
AI Data Center Operations
- Provide overall leadership for facilities operations supporting high-density AI/GPU compute environments.
- Ensure the facility consistently delivers the required power, cooling, environmental conditions, and infrastructure availability for AI compute clusters.
- Establish operational readiness standards for new AI compute deployments, including infrastructure validation, capacity verification, and turnover.
- Lead facility response to electrical, mechanical, thermal, controls, and cooling-related incidents.
High-Density Power Infrastructure
- Oversee utility service, substations, medium-voltage distribution, transformers, switchgear, UPS systems, generators, busways, PDUs, and rack-level power distribution.
- Partner with utilities and engineering teams on utility capacity, grid interconnection, power availability, resiliency, and future expansion.
- Evaluate power architectures and redundancy strategies appropriate for AI infrastructure.
- Develop long-range power capacity plans aligned with GPU/accelerator deployment schedules.
Advanced Cooling & Thermal Management
- Lead the design, operation, and optimization of cooling systems supporting high-density AI compute.
- Establish standards for liquid-cooling reliability, leak detection, water quality, filtration, redundancy, and maintenance.
- Partner with AI hardware and compute teams to understand evolving thermal requirements for GPUs, accelerators, CPUs, networking equipment, and future-generation platforms.
- Evaluate cooling capacity at the rack, row, room, and facility level.
- Develop strategies for increasing rack density without compromising thermal performance or reliability.
AI Capacity Planning & Infrastructure Strategy
- Develop multi-year facilities capacity plans based on AI compute roadmaps and anticipated GPU/accelerator deployments.
- Translate compute requirements into MW, cooling tonnage, rack density, water, space, and infrastructure requirements.
- Partner with AI infrastructure and data center engineering teams to establish deployment timelines and facility readiness requirements.
- Develop scalable architectures capable of supporting rapid expansion from individual clusters to multi-megawatt AI campuses.
- Evaluate emerging infrastructure technologies and determine their suitability for large-scale AI environments.
- Participate in site selection, utility strategy, infrastructure due diligence, and campus master planning.
New Construction, Expansion & Commissioning
- Lead facilities participation in the design, construction, commissioning, and turnover of new AI data center capacity.
- Establish commissioning requirements for electrical, mechanical, controls, liquid-cooling, and life-safety systems.
- Review engineering designs, equipment selections, redundancy models, sequence-of-operations documents, and commissioning plans.
- Develop standardized infrastructure designs that can be replicated across multiple AI data center sites.
Reliability & Incident Management
- Establish a world-class reliability program for AI data center infrastructure.
- Lead root-cause analysis for critical facility incidents and implement sustainable corrective actions.
- Develop predictive and condition-based maintenance strategies for critical assets.
- Establish reliability metrics for power, cooling, controls, and liquid-cooling systems.
- Conduct failure-mode analysis and identify potential infrastructure risks before they become operational events.
- Lead emergency response and recovery for facility incidents affecting AI compute availability.
- Conduct regular scenario-based exercises involving major power failures, cooling failures, loss of utility service, controls failures, and liquid-cooling events.
Automation, Controls & Data
- Drive increased use of BMS, EPMS, DCIM, telemetry, analytics, and predictive monitoring across the facility.
- Establish real-time visibility into electrical and thermal capacity.
- Identify opportunities to automate facility operations while maintaining appropriate safeguards and human oversight.
Team Leadership
- Build and lead a high-performing organization of facilities engineers, managers, technicians, and specialized contractors.
- Establish organizational structures capable of supporting 24x7 AI data center operations at scale.
- Develop technical training programs focused on high-density power, liquid cooling, controls, and AI infrastructure.
- Establish succession planning and technical career-development programs.
- Foster a culture of operational excellence, safety, urgency, accountability, and continuous improvement.
Financial & Vendor Management
- Manage strategic relationships with utilities, OEMs, engineering firms, contractors, cooling providers, and equipment manufacturers.
- Establish performance requirements and service-level agreements for critical vendors.
- Evaluate lifecycle costs, reliability, availability, maintainability, and scalability when selecting infrastructure technologies.
- Identify opportunities to reduce operating costs while maintaining or improving infrastructure reliability.
Safety, Compliance & Risk
- Establish a safety-first culture appropriate for high-voltage electrical systems, heavy mechanical equipment, industrial cooling systems, and construction environments.
- Develop comprehensive emergency response and business-continuity programs.
- Maintain accurate facility documentation, drawings, asset records, operating procedures, and maintenance histories.
- Conduct regular infrastructure risk assessments and resilience reviews.
- Ensure facilities are prepared for internal, customer, regulatory, insurance, and third-party audits.
Energy & Sustainability
- Develop strategies to manage the significant energy demands associated with AI compute.
- Optimize PUE, WUE, power utilization, cooling efficiency, and carbon impact.
- Partner with utilities and energy teams on renewable energy, energy procurement, demand management, and grid-related initiatives.
- Evaluate alternative cooling technologies, heat-reuse opportunities, water-conservation strategies, and other sustainability initiatives.
- Balance sustainability objectives with AI compute availability, reliability, and growth requirements.
Requirements
- Bachelor’s degree in Electrical Engineering, Mechanical Engineering, Facilities Engineering, or a related technical discipline; equivalent experience will be considered.
- 10+ years of experience in mission-critical data centers, high-density computing environments, critical infrastructure, or comparable facilities.
- 5+ years of experience leading large technical facilities organizations.
- Demonstrated experience managing large-scale electrical and mechanical infrastructure.
- Strong understanding of high-density data center power architectures and cooling systems.
- Experience supporting or operating facilities with high rack power densities and rapidly scaling compute infrastructure.
- Experience with mission-critical maintenance, reliability engineering, incident management, and operational readiness.
- Demonstrated experience managing significant operating and capital budgets.
- Strong understanding of infrastructure capacity planning and facility expansion.
- Excellent executive
Benefits & conditions
- Stock Options
- 100% paid Medical, Dental, and Vision insurance for Employees
- Company Health Savings Account Contributions
- 100% paid Short Term and Long Term Disability Insurance for Employees
- Life and Voluntary Supplemental Insurance Options
- Other Insurance Options, such as Pet & Legal Insurance
- Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
- Flexible Spending Account
- 401(k)
- Employee Assistance Program
- Flexible PTO
- Paid Holidays
- Parental Leave
- Other In-Office Perks
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
Stephan Gillich - Bringing AI Everywhere
What Industries Outside of AI Are Hiring The Most AI Experts?
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production