> Markdown version of [/jobs/ext/2718345-data-center-facilities-operations-lead](https://www.wearedevelopers.com/jobs/ext/2718345-data-center-facilities-operations-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Center Facilities Operations Lead - **Company:** Gimlet Labs Inc. - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Computer Clusters, Data Centers, Data Center Infrastructure Management (CIM), PL-SQL, AI Infrastructure, High Performance Computing - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/data-center-facilities-operations-lead-gimlet-labs-8995784 ## About the Role * You have experience in data center facilities operations, critical facilities engineering, MEP operations, commissioning, facilities maintenance, or high-density infrastructure operations. * You understand liquid cooling operations, facility water systems, CDUs, heat rejection, supply and return temperature management, flow, pressure, filtration, leak detection, and water quality controls. * You have operated or supported BMS, DCIM, EPMS, CDU monitoring, alarm response, trend analysis, and facilities telemetry in production environments. * You can write and run MOPs, SOPs, EOPs, maintenance plans, incident procedures, and vendor repair workflows with strong operational discipline. * You communicate clearly with site teams, network teams, deployment TPMs, engineering, colocation providers, OEMs, and facilities vendors. * You are comfortable working in active data center environments and supporting urgent facilities escalations when they arise. Strong candidates may also have * Experience supporting GPU clusters, GB200/GB300-class platforms, NVL rack-scale systems, HPC environments, AI infrastructure, or other high-density liquid-cooled compute deployments. * Experience with data center commissioning, integrated systems testing, site acceptance testing, facility turnover, or deployment readiness reviews. * Familiarity with power distribution, UPS/generator coordination, chilled water systems, dry coolers, CRAH/CRAC systems, CDUs, rear-door heat exchangers, and liquid cooling safety practices. * Experience managing colocation provider obligations, service levels, maintenance windows, vendor escalations, and facilities contract deliverables. * A track record of improving facility reliability through monitoring, preventive maintenance, incident analysis, documentation, training, and operational controls. ## Description Gimlet Labs is seeking a Data Center Facilities Operations Lead to own the critical facilities operating model for Gimlet data centers and high-density AI infrastructure deployments. In this role, you will make sure the facility-side systems that support Gimlet's compute capacity are ready, monitored, maintained, and operating inside the required envelope. You will focus on the infrastructure that keeps liquid-cooled AI systems healthy: facility water loops, CDUs, supply and return temperatures, flow, pressure, water quality, leak detection, alarms, heat rejection, power and cooling coordination, BMS/DCIM telemetry, maintenance procedures, and vendor repair workflows. This role is well-suited for a critical facilities operator who understands data center MEP systems, liquid cooling, operational monitoring, and the discipline required to keep high-density compute environments stable as Gimlet scales. What success looks like In the first 12-18 months, you will: * Build the facilities operations model for current and future Gimlet sites, including operating standards, escalation paths, maintenance routines, acceptance criteria, and facility readiness gates. * Translate OEM and engineering requirements for liquid-cooled platforms into practical site operating envelopes for temperature, flow, pressure, water quality, alarms, and heat rejection. * Own monitoring and response for facility-side telemetry, including supply and return water temperatures, delta-T, flow, pressure, leak detection, CDU status, cooling capacity margins, and BMS/DCIM alarms. * Partner with colocation providers, facility vendors, OEMs, Site Managers, Data Center Technicians, Deployment Leads, and TPMs to ensure facilities are ready before new compute capacity is deployed. * Create and maintain MOPs, SOPs, EOPs, maintenance windows, runbooks, inspection routines, and incident response procedures for critical facilities and liquid cooling operations. * Coordinate preventive maintenance, repairs, and vendor response for CDUs, facility water loops, filters, valves, pumps, sensors, leak detection systems, chillers, dry coolers, CRAHs, and related infrastructure. * Lead facility-side root cause analysis for thermal, leak, power, cooling, monitoring, and environmental events, then drive durable corrective actions. * Build reporting that shows facility health, risk, readiness, capacity margin, recurring issues, open repairs, and operational trends across Gimlet sites. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [How to develop an autonomous car end-to-end: Robotic Drive and the mobility revolution](https://www.wearedevelopers.com/videos/22-how-to-develop-an-autonomous-car-end-to-end-robotic-drive-and-the-mobility-revolution) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [5 Years in Cloud Native: The Good, the Bad, and the Bill](https://www.wearedevelopers.com/videos/100111-5-years-in-cloud-native-the-good-the-bad-and-the-bill) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)