> Markdown version of [/jobs/ext/2941966-manager-data-center-facilities-engineering](https://www.wearedevelopers.com/jobs/ext/2941966-manager-data-center-facilities-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Manager, Data Center Facilities Engineering - **Company:** Nebius Inc. - **Location:** Philadelphia, United States (Remote available) - **Experience:** Expert - **Salary:** $115,000.0 - $275,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computerized Maintenance Management Systems, Data Centers, Data Center Infrastructure Management (CIM), Monitoring of Systems, Machine Learning, System Availability, Information Technology, Performance Monitor - **Published:** September 16, 2026 - **Apply:** https://startup.jobs/manager-data-center-facilities-nebius-8329981 ## About the Role * Bachelor's degree in Engineering, Mechanical/Electrical Technology, or a related technical field-or equivalent practical experience * 10+ years of experience in critical facility or data center operations, or other mission-critical environments * 5+ years of experience leading and developing technical teams, including performance management * Strong understanding of critical infrastructure systems: electrical distribution (UPS, generators, switchgear), mechanical/HVAC, fire/life safety, and building automation/controls * Experience operating in 24/7 high-availability environments with strict uptime requirements * Working knowledge of procedure-based operations and safety programs (e.g., LOTO, electrical safety, hazardous energy control) * Proven ability to collaborate with cross-functional teams including construction, engineering, and operations * Strong communication skills with experience presenting to executive leadership and creating clear, concise reports and documentation * Experience managing vendors, maintenance programs, and incident response * Familiarity with SOPs, EOPs, MOPs, and operational monitoring systems (BMS/DCIM) It would be an added bonus if you have: * Experience in data center or Tier III/IV critical environments, preferably within hyperscale or high-availability operations * Advanced education (Master's or MBA) or relevant trade certification (Electrical, HVAC, Controls) * Experience supporting commissioning, new builds, or large-scale facility expansions * Strong knowledge of critical infrastructure systems, including power, cooling (CRAH/CRAC, chillers), and exposure to modern technologies such as liquid cooling * Proven ability to drive operational excellence through budget management, continuous improvement (Lean/Six Sigma), vendor oversight, and use of tools like CMMS/EAM and AI-enabled workflows * Experience operating in hyperscale, AI/ML-driven environments supporting GPU-intensive, high-density workloads * Proven ability to scale data center operations to meet rapid growth and next-generation infrastructure demands * Strong focus on automation, telemetry, and data-driven decision-making to optimize performance and reliability * Experience driving efficiency in advanced cooling environments, including liquid cooling and high-density thermal management, Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. ## Description The Manager, Data Center Facilities Engineering is a critical leadership role responsible for ensuring 24/7/365 availability of data center operations by overseeing maintenance and repair of critical infrastructure in alignment with SLAs. This role leads emergency response, coordinates cross-functional teams, and ensures seamless communication and compliance with operational and safety standards. The manager provides oversight of contracts, maintenance plans, and policy development while driving performance through metrics, continuous improvement, and team development. Your responsibilities will include: * Own day-to-day operations of critical electrical and mechanical systems, including UPS, generators, switchgear, PDUs, chillers, and CRAH/CRAC units * Ensure high availability and uptime of all facility infrastructure supporting data center operations * Lead incident response for facility-related events, including root cause analysis and implementation of corrective actions * Monitor performance through BMS/DCIM systems and drive improvements in reliability, efficiency, and capacity utilization * Oversee preventive and corrective maintenance programs, ensuring adherence to SOPs, EOPs, and MOPs * Manage and hold third-party vendors accountable for service delivery, performance, and compliance with standards * Partner with engineering and operations teams on capacity planning, infrastructure scaling, and site optimization initiatives * Drive energy efficiency and sustainability efforts, including PUE optimization and support for high-density cooling solutions * Ensure compliance with safety, regulatory, and HSE standards, leading audits, inspections, and risk management activities * Collaborate cross-functionally with data center operations, network, and infrastructure teams to ensure seamless integration between facilities and IT systems ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Autonomous microservices with event-driven architecture](https://www.wearedevelopers.com/videos/1052-autonomous-microservices-with-event-driven-architecture) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [How to build a sovereign European AI compute infrastructure](https://www.wearedevelopers.com/videos/1102-how-to-build-a-sovereign-european-ai-compute-infrastructure) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [6 Emerging Technologies We’ll Learn About in 2025](https://www.wearedevelopers.com/magazine/381-6-emerging-technologies-we-ll-learn-about-in-2025) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)