Project Manager (Major Incident Management & NOC )
American IT Systems
Wilmington, United States of America
2 days ago
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
English Experience level
SeniorJob location
Wilmington, United States of America
Tech stack
Microsoft Windows
Amazon Web Services (AWS)
Azure
Cloud Computing
Configuration Management Databases
Databases
Dynamic Host Configuration Protocol
Noise Reduction
DNS
Hyper-V
Identity and Access Management
Linux System Administration
Networking Basics
Routing
Oracle Applications
Performance Tuning
SQL Databases
TCP/IP
Datadog
Data Logging
Google Cloud Platform
Load Balancing
Mttr
Firewalls (Computer Science)
Amazon Web Services (AWS)
Cloudwatch
Splunk
New Relic (SaaS)
Appdynamics
Dynatrace
ServiceNow
VMware
Job description
- A) Manage Team of Major Incident Managers (Command & Control)
- Own the Major Incident (P1/P2) process from detection to resolution, including war-room leadership, stakeholder updates, and closure.
- Ensure structured triage, containment, workaround, and restoration.
- Drive cross-functional coordination (App, Infra, Network, Security, DB, Cloud, Vendor teams) to reduce MTTR.
- Ensure high-quality incident communications: executive summaries, impact analysis, ETAs, customer/business comms.
- Lead and facilitate Post Incident Reviews (PIR/RCA); ensure actionable corrective/preventive actions (CAPA).
- Identify recurring issues and trigger Problem Management with measurable reduction plans.
- ITIL v4 Foundation (preferred).
- Reduced MTTD and MTTR for P1/P2 incidents.
- Improved SLA compliance and reduction in escalation breaches.
- Reduced repeat incidents via problem management and preventive actions.
- Improved alert quality: lower false positives, better signal-to-noise ratio.
- Strong PIR/RCA compliance: on-time RCAs with measurable preventive outcomes.
- Improved NOC operational maturity: SOP adherence, shift handover quality, audit readiness.
- B) NOC Leadership & Operations
- Manage the NOC team responsible for 24x7 monitoring, alert triage, event correlation, escalation, and ticket quality.
- Establish/maintain standard operating procedures (SOPs), runbooks, escalation matrices, and on-call models.
- Ensure NOC meets SLAs/OLAs, improves alert fidelity, and reduces noise through tuning and automation.
- Manage handover governance between shifts; maintain service continuity and operational hygiene.
- C) Service Reliability & Continuous Improvement
- Drive operational improvements: monitoring coverage, SLO/SLA alignment, incident prevention, and resiliency initiatives.
- Partner with Engineering/Platform teams on observability strategy, proactive detection, and reliability patterns.
- Track and report operational metrics: MTTD, MTTR, incident volume, re-open rate, SLA compliance, and trends.
- Support readiness for audits and compliance: evidence collection, process adherence, and risk mitigation.
- D) Stakeholder & Vendor Management
- Interface with business stakeholders, service owners, and leadership to provide incident status, risk, and remediation plans.
- Manage vendor escalations and ensure timely resolution aligned to contractual SLAs.
Requirements
- Proven experience leading MIM & NOC Operations teams (shift-based or on-call models).
- Excellent stakeholder management across technical teams and business leadership.
- Ability to build and enforce process discipline (ITIL-aligned), while improving speed and quality.
- Strong coaching/mentoring: performance management, skill development, hiring support as needed.
- Effective communication: concise executive updates, clear action plans, facilitation of PIR/RCA sessions.
- Data-driven mindset: uses metrics and trend analysis to drive operational outcomes.
Technical Skills (Must Have):
- A) Monitoring / Observability
Strong understanding of event correlation, alert tuning, noise reduction, and dashboarding.
Strong working knowledge of ServiceNow (Incident, Problem, Change, Knowledge, CMDB) or equivalent ITSM tools.
- B) Incident / ITSM Platforms (Good to Have )
Experience designing workflows, SLAs/OLAs, routing rules, and automation integrations.
Hands-on experience with NOC tooling and observability platforms such as:Splunk / ELK, Datadog, Dynatrace, New Relic, AppDynamics,PrometheGrafana, CloudWatch/Azure Monitor
- C) Infrastructure & Platform Breadth
- Solid understanding across:
- Windows/Linux administration basics
- Network fundamentals (DNS, DHCP, TCP/IP, routing, load balancers, firewalls)
- Compute/virtualization (VMware/Hyper-V) and storage concepts
- Databases fundamentals (SQL/Oracle, replication, performance symptoms)
- Cloud fundamentals and operational support for AWS/Azure/Google Cloud Platform:
- IAM basics, networking (VPC/VNet), scaling, logging/monitoring, common failure patterns.