Data centre Site Operations Leader

Digipowers, Inc.
Columbiana, AL, United States
25 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$124,800.0 - $208,000.0
Working hours
Regular working hours
Job source

Tech stack

Applicant Tracking Systems Data Centers Data Center Infrastructure Management (CIM) Monitoring of Systems AI Infrastructure Information Technology Hardware Infrastructure 3-tier Architectures

Job description

DigiPowerX operates AI infrastructure facilities serving GPU-as-a-Service and colocation customers. We are hiring a Site Operations Leader to own day-to-day operations at our Alabama facility and will be accountable for operational excellence as we bring additional facilities online. Your goals are to operate this site at - Tier 3 (99.982 uptime), GPU infrastructure (99.9 availability), 100% SOP/MOP coverage & Change Management.

This is a hands-on, technical leadership role, based at our site. The right person has operated critical data center facilities, CPU/GPU infrastructures, built maintenance programs, managed onsite teams, and held vendors accountable. You will be the most senior operational authority onsite.

Responsibilities

Facility & AI Infrastructure operations

  • Own day-to-day operations of live, GPU-dense, liquid-cooled data center facilities* Build and maintain a structured preventative maintenance program for all MEP systems: chiller plant, generators, ATS, fire suppression, and power distribution* Develop and maintain SOPs, MOPs, and emergency operating procedures grounded in industry best practices* Serve as onsite incident commander for MEP and infrastructure events* Monitor facility health through BMS/DCIM systems and drive resolution of anomalies before they become incidents* Monitor CPU/GPU rack health, and Network health through DCIM systems and drive resolution of anomalies before they become incidents. Drive resolutions and escalations with support vendors. * Manage RMA process with support vendors of GPU, MEP components while holding them accountable to their contracted SLA/SLO

Team leadership

  • Manage a growing team of multicraft technicians and site support personnel* Manage a team of smart hands operators with skills in networking, racks and compute infrastructure* Own work assignment, scheduling, shift coverage, performance management, and training* Build a metrics driven culture of operational discipline, accountability, and continuous improvement

Vendor and MSP management

  • Interface with MEP vendors, ensuring contractual obligations and SLAs are met* Provide oversight of managed service providers supporting current and future facilities* Hold vendors accountable through documented performance tracking and regular business reviews

Compliance and commissioning

  • Support ISO 27001, ISO 22237, and SOC 2 certification efforts by owning facility-level evidence generation and operational documentation* Provide operational oversight during commissioning of new facilities, working alongside commissioning engineers and compliance partners with decision authority

Requirements

  • 5+ years of experience operating mission-critical customer facing data center facilities (not IT infrastructure or enterprise server rooms)* Direct experience building or managing a preventative maintenance program for MEP systems (cooling, power, fire suppression)* Exposure to compliance programs such as SOC 2, ISO 27001, or ISO 22237.* Demonstrated understanding of chiller plant operations, generator/ATS testing, and fire suppression systems* Experience managing onsite technical teams in a 24/7 critical environment* Experience with vendor and/or MSP oversight and accountability* Familiarity with BMS/DCIM monitoring systems* Willingness to be based onsite in Columbiana, Alabama

Preferred

  • Experience with liquid-cooled, high-density GPU compute environments* Experience with facility commissioning from the operator’s perspective* Background in hyperscale, colocation, or AI infrastructure environments

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:53 min

Evolution of computing to AI infrastructure

Michael Kagan Michael Kagan +1 · WWC Europe 2026

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:03 min

Optimizing energy consumption and sustainability in data centers

Markus Hacker Markus Hacker +3 · WWC 2024

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · WWC 2024

2:09 min

Treating enterprise AI configurations as customizable internal infrastructure

Thomas Froment Thomas Froment · WWC Europe 2026

Videos

See all

Related articles

See all