Manager Major Incident & Problem Management

LTD Global
Gainesville, FL, United States
3 days ago
Apply on jobs.localjobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$95,000.0 - $108,000.0
Working hours
Regular working hours

Tech stack

Configuration Management Knowledge Management Information Technology Servicenow

Job description

We are seeking a Manager to own the execution and continued maturity of two critical IT Service Management practices, Major Incident & Problem Management. This role leads the response to all enterprise Priority 1 incidents, coordinates rapid restoration of business services, and ensures leaders and stakeholders receive timely, accurate, and business-focused communications.

Beyond active incident response, this position sets the strategic direction for Major Incident and Problem Management. The role provides functional, dotted-line leadership to the existing MSP (Managed Service Provider) Outage Coordinators and Problem Analyst, establishes consistent operating standards, drives accountability for root-cause and corrective-action work, and uses operational insights to reduce repeat incidents and improve service reliability.

This is a hands-on leadership role for someone who can remain composed during high-impact events, bring structure to ambiguity, influence teams without relying on direct authority, and translate technical conditions into clear business impact and decisions. What You’ll DoEnterprise Major Incident Leadership

  • Own and lead the end-to-end response for all enterprise P1 incidents, from declaration and bridge activation through service restoration, stakeholder transition, and formal closure.
  • Establish command and control during major incidents by clarifying roles, driving urgency, maintaining decision discipline, and ensuring the right technical and business resources are engaged.
  • Facilitate incident bridges, maintain focus on restoration, remove coordination obstacles, and escalate risks or resource gaps to technology leadership.
  • Ensure business impact, scope, workarounds, recovery progress, and restoration status are validated before they are communicated.
  • Coordinate executive, technology, and business communications using clear, concise, and audience-appropriate messaging.
  • Lead post-incident reviews and confirm that key decisions, timelines, lessons learned, and follow-up actions are documented.

Problem Management Ownership

  • Own the enterprise Problem Management practice, including intake, prioritization, investigation governance, known-error discipline, and closure criteria.
  • Ensure significant and recurring incidents are evaluated for problem records and that root-cause analysis is completed with appropriate rigor.
  • Drive accountable corrective and preventive actions with named owners, target dates, evidence of completion, and risk-based escalation for overdue work.
  • Partner with engineering, infrastructure, application, vendor, and service-owner teams to eliminate systemic causes and reduce recurrence.
  • Identify patterns across incidents, problems, changes, monitoring events, and service dependencies to inform reliability priorities.

Functional Team Leadership

  • Provide dotted-line leadership, operating direction, coaching, and quality oversight for MSP provided Outage Coordinators and the Problem Analyst.
  • Define role expectations, coverage models, escalation paths, facilitation standards, documentation requirements, and communication quality controls.
  • Conduct case reviews and targeted coaching to build consistency, confidence, and sound judgment across the team.
  • Coordinate workload and coverage with internal leaders and vendor management while maintaining clear accountability for practice outcomes.
  • Serve as the escalation point for complex incidents, stalled investigations, unresolved ownership, and process exceptions.

Strategy, Governance & Continuous Improvement

  • Develop and maintain the multi-year strategy, roadmap, operating model, policies, procedures, playbooks, and maturity plan for Major Incident and Problem Management.
  • Establish governance forums and performance reviews that focus on outcomes, risks, recurring failure themes, corrective-action health, and improvement priorities.
  • Define and monitor meaningful measures such as restoration performance, communication timeliness and quality, recurrence, root-cause completion, action aging, and business impact.
  • Identify opportunities to automate workflows, notifications, evidence capture, reporting, and handoffs across ITSM systems and adjacent platforms.
  • Align the practices with IT Service Management standards and integrate them with Change, Configuration, Knowledge, Event, Service Level, and Continuity Management.
  • Create training and simulation exercises that strengthen incident leadership, technical response, business-impact assessment, and executive communication.

Requirements

  • 5+ years of progressive experience in IT Service Management, service operations, incident management, problem management, or a related enterprise technology function.
  • Demonstrated experience leading high-severity incidents in a complex, multi-team environment with material business impact.
  • Experience designing, maturing, or governing Major Incident and Problem Management processes, not only executing individual cases.
  • Experience leading internal teams, managed-service providers, or matrixed resources through influence and clearly defined accountability.
  • Experience presenting incident status, risk, root cause, and corrective-action progress to senior technology and business leaders.

Operational & Technical Depth

  • Strong working knowledge of ITIL practices, particularly Incident Management, Major Incident Management, Problem Management, Change Enablement, Configuration Management, Knowledge Management, and Service Level Management.
  • Practical experience with an enterprise ITSM platform; ServiceNow experience is strongly preferred.
  • Ability to understand complex application, infrastructure, network, cloud, integration, and vendor dependencies sufficiently to lead restoration and challenge assumptions.
  • Ability to use incident and problem data to identify trends, quantify operational risk, and prioritize improvement opportunities.
  • Comfort with on-call or after-hours engagement when enterprise P1 incidents require leadership.

Leadership & Communication

  • Calm, decisive, and highly organized during fast-moving, high-pressure events.
  • Exceptional facilitation skills with the ability to maintain urgency without creating noise or confusion.
  • Clear writer and communicator who can translate technical detail into business impact, decisions, risks, and next steps.
  • Strong judgment, ownership, follow-through, and willingness to escalate when service restoration or corrective action is at risk.
  • Collaborative and credible with technical teams, business stakeholders, executives, and external partners.

Education

Bachelor’s degree in Information Technology, Computer Science, Business, or a related field, or equivalent practical experience. Preferred Certifications

  • ITIL 4 or 5 Foundation; ITIL Practice Manager, Monitor, Support and Fulfil, or equivalent advanced ITSM certification.
  • ServiceNow Certified System Administrator, Certified Implementation Specialist - IT Service Management, or equivalent platform experience.
  • Relevant incident command, problem analysis, reliability, or project leadership certification.

About the company

Why Choose GMR? Global Medical Response(GMR) and its family of solutions are dedicated to delivering compassionate, quality medical care, primarily in the areas of emergency and patient relocation services. Here you’ll embark in meaningful work that will make an impact on you and the customers we service. View our employees’ stories on how we provide care to the world at www.AtaMomentsNotice.com.

GMR’s Core Behaviors-keep care at the center, raise your hand,seekto understand, find a way together and be accountable-uniteour teamsand set us apart in emergency medical services., Global Medical Response and its family of companies are an Equal Opportunity Employer, which includes supporting veterans and providing reasonable accommodations for individuals with a disability.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.localjobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

56 sec

Integrating automated approval workflows into the portal

Markus Eisele Markus Eisele · World Congress 2025

2:52 min

Scaling generative AI use cases across large enterprises

Mike Butcher Mike Butcher +3 · World Congress 2024

42 sec

Energy forecasts and resource demands of information technology

Marjolein Pordon · LIVE

4:05 min

Maximizing global incident coverage through asynchronous remote team distribution

Hazal Mestci +1 · Coffee With Developers

2:27 min

Establishing a simulated technical environment for the workflow demo

Tobias Dunn-Krahn · LIVE

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

Videos

See all

Related articles

See all