> Markdown version of [/jobs/ext/3569408-manager-major-incident-problem-management](https://www.wearedevelopers.com/jobs/ext/3569408-manager-major-incident-problem-management). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Manager Major Incident & Problem Management - **Company:** LTD Global - **Location:** Gainesville, FL, United States (Remote available) - **Experience:** Expert - **Salary:** $95,000.0 - $108,000.0 - **Contract:** Permanent contract - **Skills:** Configuration Management, Knowledge Management, Information Technology, Servicenow - **Published:** October 3, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88508328/1 ## About the Role * 5+ years of progressive experience in IT Service Management, service operations, incident management, problem management, or a related enterprise technology function. * Demonstrated experience leading high-severity incidents in a complex, multi-team environment with material business impact. * Experience designing, maturing, or governing Major Incident and Problem Management processes, not only executing individual cases. * Experience leading internal teams, managed-service providers, or matrixed resources through influence and clearly defined accountability. * Experience presenting incident status, risk, root cause, and corrective-action progress to senior technology and business leaders. Operational & Technical Depth * Strong working knowledge of ITIL practices, particularly Incident Management, Major Incident Management, Problem Management, Change Enablement, Configuration Management, Knowledge Management, and Service Level Management. * Practical experience with an enterprise ITSM platform; ServiceNow experience is strongly preferred. * Ability to understand complex application, infrastructure, network, cloud, integration, and vendor dependencies sufficiently to lead restoration and challenge assumptions. * Ability to use incident and problem data to identify trends, quantify operational risk, and prioritize improvement opportunities. * Comfort with on-call or after-hours engagement when enterprise P1 incidents require leadership. Leadership & Communication * Calm, decisive, and highly organized during fast-moving, high-pressure events. * Exceptional facilitation skills with the ability to maintain urgency without creating noise or confusion. * Clear writer and communicator who can translate technical detail into business impact, decisions, risks, and next steps. * Strong judgment, ownership, follow-through, and willingness to escalate when service restoration or corrective action is at risk. * Collaborative and credible with technical teams, business stakeholders, executives, and external partners. Education Bachelor's degree in Information Technology, Computer Science, Business, or a related field, or equivalent practical experience. Preferred Certifications * ITIL 4 or 5 Foundation; ITIL Practice Manager, Monitor, Support and Fulfil, or equivalent advanced ITSM certification. * ServiceNow Certified System Administrator, Certified Implementation Specialist - IT Service Management, or equivalent platform experience. * Relevant incident command, problem analysis, reliability, or project leadership certification. ## Description We are seeking a Manager to own the execution and continued maturity of two critical IT Service Management practices, Major Incident & Problem Management. This role leads the response to all enterprise Priority 1 incidents, coordinates rapid restoration of business services, and ensures leaders and stakeholders receive timely, accurate, and business-focused communications. Beyond active incident response, this position sets the strategic direction for Major Incident and Problem Management. The role provides functional, dotted-line leadership to the existing MSP (Managed Service Provider) Outage Coordinators and Problem Analyst, establishes consistent operating standards, drives accountability for root-cause and corrective-action work, and uses operational insights to reduce repeat incidents and improve service reliability. This is a hands-on leadership role for someone who can remain composed during high-impact events, bring structure to ambiguity, influence teams without relying on direct authority, and translate technical conditions into clear business impact and decisions. What You'll DoEnterprise Major Incident Leadership * Own and lead the end-to-end response for all enterprise P1 incidents, from declaration and bridge activation through service restoration, stakeholder transition, and formal closure. * Establish command and control during major incidents by clarifying roles, driving urgency, maintaining decision discipline, and ensuring the right technical and business resources are engaged. * Facilitate incident bridges, maintain focus on restoration, remove coordination obstacles, and escalate risks or resource gaps to technology leadership. * Ensure business impact, scope, workarounds, recovery progress, and restoration status are validated before they are communicated. * Coordinate executive, technology, and business communications using clear, concise, and audience-appropriate messaging. * Lead post-incident reviews and confirm that key decisions, timelines, lessons learned, and follow-up actions are documented. Problem Management Ownership * Own the enterprise Problem Management practice, including intake, prioritization, investigation governance, known-error discipline, and closure criteria. * Ensure significant and recurring incidents are evaluated for problem records and that root-cause analysis is completed with appropriate rigor. * Drive accountable corrective and preventive actions with named owners, target dates, evidence of completion, and risk-based escalation for overdue work. * Partner with engineering, infrastructure, application, vendor, and service-owner teams to eliminate systemic causes and reduce recurrence. * Identify patterns across incidents, problems, changes, monitoring events, and service dependencies to inform reliability priorities. Functional Team Leadership * Provide dotted-line leadership, operating direction, coaching, and quality oversight for MSP provided Outage Coordinators and the Problem Analyst. * Define role expectations, coverage models, escalation paths, facilitation standards, documentation requirements, and communication quality controls. * Conduct case reviews and targeted coaching to build consistency, confidence, and sound judgment across the team. * Coordinate workload and coverage with internal leaders and vendor management while maintaining clear accountability for practice outcomes. * Serve as the escalation point for complex incidents, stalled investigations, unresolved ownership, and process exceptions. Strategy, Governance & Continuous Improvement * Develop and maintain the multi-year strategy, roadmap, operating model, policies, procedures, playbooks, and maturity plan for Major Incident and Problem Management. * Establish governance forums and performance reviews that focus on outcomes, risks, recurring failure themes, corrective-action health, and improvement priorities. * Define and monitor meaningful measures such as restoration performance, communication timeliness and quality, recurrence, root-cause completion, action aging, and business impact. * Identify opportunities to automate workflows, notifications, evidence capture, reporting, and handoffs across ITSM systems and adjacent platforms. * Align the practices with IT Service Management standards and integrate them with Change, Configuration, Knowledge, Event, Service Level, and Continuity Management. * Create training and simulation exercises that strengthen incident leadership, technical response, business-impact assessment, and executive communication. ## Related Videos - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Enabling intelligent logistics automation: home-grown Industrial IoT platform at Austrian Post](https://www.wearedevelopers.com/videos/2018-enabling-intelligent-logistics-automation-home-grown-industrial-iot-platform-at-austrian-post) - [AI or KO: Is HR ever going to use intelligent technology?](https://www.wearedevelopers.com/videos/1473-ai-or-ko-is-hr-ever-going-to-use-intelligent-technology) - [It's not easy being green](https://www.wearedevelopers.com/videos/558-it-s-not-easy-being-green) - [Same Tower, New Confusion: The Tower of Babel 2.0](https://www.wearedevelopers.com/videos/100101-same-tower-new-confusion-the-tower-of-babel-2-0) - [AI in Production: applied AI & enterprise use cases](https://www.wearedevelopers.com/videos/100130-ai-in-production-applied-ai-enterprise-use-cases) ## Related Articles - [From developer to manager – what does it take to become an engineering manager?](https://www.wearedevelopers.com/magazine/42-from-developer-to-manager-what-does-it-take-to-become-an-engineering-manager) - [The Geometry of Incidents: Connecting User Impact to Architecture](https://www.wearedevelopers.com/magazine/764-the-geometry-of-incidents-connecting-user-impact-to-architecture) - [Should Tech Managers Be Developers First? Pros and Cons](https://www.wearedevelopers.com/magazine/327-should-tech-managers-be-developers-first-pros-and-cons) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023) - [Promotion Interview Questions: How to Answer and Get the Job](https://www.wearedevelopers.com/magazine/413-promotion-interview-questions-how-to-answer-and-get-the-job)