> Markdown version of [/jobs/ext/1434838-datacenter-operations-manager](https://www.wearedevelopers.com/jobs/ext/1434838-datacenter-operations-manager). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Datacenter Operations Manager - **Company:** WeEngage Group | B Corp - **Location:** Newcastle upon Tyne, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Computerized Maintenance Management Systems, Data Centers, Firmware, Networking Hardware, AI Infrastructure, Network Server - **Published:** July 25, 2026 - **Apply:** https://www.apply4u.co.uk/jobs/x/41945141/ ## About the Role What we are looking forAt least five years of experience within data center operations, critical infrastructure, cloud infrastructure, or another mission-critical technical environment.Strong hands-on experience installing and supporting enterprise servers, storage, networking hardware, and structured cabling.Previous experience acting as a site lead, senior technician, shift lead, operations manager, or technical escalation point.Practical understanding of data center power, cooling, networking, rack layouts, environmental monitoring, and common infrastructure failure modes.Experience managing incidents in a structured manner, including escalation, root-cause analysis, documentation, and preventive action.Ability to work independently and make sound operational decisions without requiring constant supervision.Experience coordinating technicians, contractors, vendors, and remote engineering teams.Strong planning and organisational skills, particularly around installations, maintenance windows, upgrades, and capacity expansion.Clear written and verbal communication skills.Willingness to work on-site full-time in Newcastle and participate in on-call coverage. Particularly relevant experienceSupporting enterprise GPU platforms using NVIDIA Ampere, Hopper, Blackwell, GB200, GB300, or similar systems.Operating high-density AI, HPC, hyperscale, or cloud infrastructure.Direct liquid cooling, coolant distribution units, liquid loops, or other advanced cooling technologies.Large GPU cluster deployments, server bring-up, burn-in, firmware updates, and hardware validation.Data center build-outs, new data hall openings, migrations, expansions, or infrastructure refresh programmes.DCIM, CMMS, monitoring, alerting, ticketing, and maintenance-management platforms.Managing 24/7 shift coverage or supporting teams operating across multiple shifts.Working within environments governed by strict SLAs, security controls, and safety procedures.Data center qualifications such as CDCP, CDCS, or an equivalent certification. ## Description About the opportunityWe are supporting a fast-growing AI infrastructure company that designs, deploys, and operates large-scale GPU compute environments.The company is expanding its data center operations in Newcastle, UK and is looking for a hands-on Data Center Site Lead to take ownership of the site's day-to-day operation, technical reliability, and future growth of their UK bsuiness. This is not a purely managerial position. You will be expected to understand the site in detail, including its power, cooling, networking, server infrastructure, dependencies, capacity constraints, and operational risks. You will act as the senior technical presence on-site, lead other technicians and vendors, and take ownership when incidents or equipment failures occur.The environment supports demanding AI and high-performance computing workloads where uptime, response speed, and disciplined execution are critical. Key responsibilitiesTake day-to-day operational ownership of the Newcastle data center site.Act as the senior technical lead for on-site technicians, contractors, vendors, and remote-hands teams.Install, configure, troubleshoot, and maintain GPU servers, storage systems, networking equipment, cabling, and supporting infrastructure.Monitor site conditions, including power, cooling, temperature, humidity, capacity, alarms, and equipment health.Ensure the availability and reliability of the site within a 24/7 operational environment.Lead the response to hardware failures, environmental alarms, connectivity issues, and other critical incidents.Own incident reporting from initial detection through root-cause analysis, corrective action, and final closure.Track equipment downtime, identify recurring failure patterns, and introduce preventive measures.Coordinate escalations with hardware manufacturers, colocation providers, network teams, and other technical partners.Plan and schedule server installations, rack deployments, maintenance activities, upgrades, and hardware refreshes.Support data hall expansions, cluster deployments, migrations, and new capacity coming online.Maintain accurate records covering site assets, installations, incidents, maintenance work, capacity, and operational risks.Ensure all work follows the company's safety, security, access-control, and change-management procedures.Help develop site operating procedures, escalation processes, maintenance schedules, and reliability standards.Mentor junior technicians and support the recruitment and development of the on-site team as the facility grows.Provide regular updates to leadership on uptime, incidents, staffing, capacity, operational risks, and planned work.Participate in an on-call rotation and provide escalation support for critical incidents outside normal working hours. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Kubernetes Security - Challenge and Opportunity](https://www.wearedevelopers.com/videos/412-kubernetes-security-challenge-and-opportunity) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) ## Related Articles - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [6 Emerging Technologies We’ll Learn About in 2025](https://www.wearedevelopers.com/magazine/381-6-emerging-technologies-we-ll-learn-about-in-2025) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)