IT Operations Manager
Role details
Job location
Tech stack
Job description
Seeking an experienced Operations Manager to act as the operational owner & Service Manager for a business-critical platform, while also managing the Customer Service (Level 1) function. Role combines ITIL-based service management discipline, Site Reliability Engineering (SRE) principles, & people leadership to ensure high service availability, effective incident response, & continuous improvement across both customer-facing support & Back End service operations. You will have end-to-end accountability for live service operations, leading both the Service Engineering team & the Customer Service (L1) team, & owning service performance, platform reliability, operational risk, & financial stewardship. Hands-on experience managing cloud platforms in critical, always-on environments is a mandatory requirement for this role.
IT Operations Manager Key Responsibilities Service & Operational Ownership
- Act as the named Service Manager for the platform, with full accountability for service performance, stability, & customer impact.
- Own the service life cycle, from operational readiness & go-live through live service management & continual improvement.
- Define, own, & report against SLAs, SLOs, & operational KPIs across both customer service & service engineering functions.
- Serve as the primary operational escalation point for internal stakeholders & key customers.
Customer Service (L1) Management
- Lead & manage the Customer Service (Level 1) team, ensuring consistent, high-quality first-line support for customers.
- Ensure effective triage, prioritisation, & escalation of incidents from L1 to Service Engineering.
- Drive customer-focused service metrics, including response times, resolution quality, & customer satisfaction.
- Establish training, coaching, & quality assurance processes to continually improve L1 service delivery.
Reliability, Availability & Incident Management
- Own the end-to-end reliability & availability of a mission-critical, compliance focused platform.
- Apply SRE principles to reduce incidents, manage operational risk, & balance reliability with delivery velocity.
- Lead major incident management, ensuring effective coordination, clear communication, & rapid service restoration.
ITIL-Aligned Service Operations
- Lead Incident, Problem, Change, & Release Management in line with ITIL best practices.
- Plan & execute on-premises software upgrades & platform changes, ensuring controlled delivery & minimal disruption.
- Drive thorough root cause analysis (RCA) & ensure corrective actions are implemented & tracked to completion.
- Maintain audit-ready service documentation, runbooks, & operational procedures.
Cloud Platform & Engineering Collaboration
- Own operational oversight of cloud & hybrid platforms supporting critical customer services.
- Work closely with Engineering, Product, & Security teams to ensure platforms are operationally ready, resilient, observable, & secure.
- Ensure appropriate monitoring, alerting, capacity planning, & resilience controls are in place across Azure & hybrid environments.
- Champion automation & Infrastructure-as-Code to reduce operational toil & improve reliability.
Budget & Cost Management
- Own & manage the operations & service engineering budget, ensuring spend is forecast, controlled, & aligned to service outcomes.
- Manage costs related to cloud infrastructure, on-premises upgrades, tooling, licensing, & third-party services.
- Partner with Finance & Procurement to justify investment & identify cost optimisation opportunities without compromising service reliability or compliance.
Leadership & Team Development
- Lead, coach, & develop both the Customer Service (L1) & Service Engineering teams.
- Establish structured onboarding, training, & progression paths to build resilient, high performing teams.
- Foster a culture of accountability, service excellence, & continuous improvement across operations.
Requirements
Required:- Proven experience managing cloud platforms in critical, always-on production environments. Demonstrable experience owning or operating services where uptime, data integrity, & regulatory compliance are critical. Azure Certification. Operations Manager, Service Manager, or equivalent, with ownership of live services. Strong hands-on experience with ITIL service management practices, particularly Incident, Problem, Change, & Continual Improvement. Experience managing Customer Service/L1 support teams in a production environment. Working knowledge of Site Reliability Engineering (SRE) principles & operational risk management. Strong technical foundation across Azure, Windows Server, Linux (RedHat), Active Directory, networking, & Scripting (PowerShell, Bash, or Python). Experience delivering platform upgrades & managing production change in cloud & hybrid environments. Experience owning operational budgets & cost centres. Calm, structured leadership style with a strong focus on uptime, customer impact, deadlines, & service quality. A genuine commitment to training, mentoring, & building high-performing operational teams.