Director, IT Resiliency, GCC, and MIM

ScienceJobs.Org
New York, NY, United States
3 days ago
Apply on sciencejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours

Tech stack

Configuration Management Databases Data Centers Disaster Recovery Monitoring of Systems HP Systems Insight Manager IT Management Performance Monitor Servicenow

Job description

The Director of IT Resiliency, Global Command Center, and Major Incident Management (IT Resiliency, GCC, and MIM) leads employees, consultants, and a global managed service provider to deliver three critical NYU IT functions: the IT Resiliency program, Global Command Center operations, and Major Incident Management. The Director oversees the University’s IT Resiliency program for enterprise administrative and academic systems and services. This includes developing strategy and operational plans for application resiliency, aligning initiatives with NYU IT’s strategic roadmap, and advising IT leadership on program progress, priorities, and risks. The role directs Application Impact Analysis (AIA) and lifecycle resiliency planning for enterprise systems, partnering with institutional stakeholders globally. The Director also collaborates with NYU Public Safety and Risk Management to ensure effective coordination between Business Continuity and IT Resiliency processes. For Global Command Center operations, the Director defines strategy and manages the 24/7 monitoring and incident response environment supporting NYU’s global systems and New York data centers. The role oversees the managed service provider, ensuring performance targets are met, and directs the evolution of monitoring capabilities, tools, and operational processes. Responsibilities include managing Tier 1 support, improving monitoring integrity and performance, identifying operational improvements, and providing leadership with regular status and performance reporting. For Major Incident Management, the Director establishes and governs the enterprise major incident strategy and operating model within ServiceNow. The role ensures critical incidents are managed quickly and consistently in alignment with business priorities. Responsibilities include defining escalation models and communication frameworks, overseeing the full incident lifecycle-from detection through resolution and post-incident review-and ensuring real-time visibility through dashboards, war rooms, and automated workflows. The Director drives cross-functional coordination across IT, security, infrastructure, and business teams; enforces SLA adherence; analyzes trends and KPIs through ServiceNow Performance Analytics; and leads continual service improvement initiatives. The role also ensures integration with monitoring systems, CMDB integrity, regulatory compliance, and effective crisis communications. The Director designs and leads the organizational structure for IT Resiliency, GCC, and MIM, ensuring appropriate staffing across employees, contractors, and managed service providers. The role evaluates team performance, drives operational excellence globally, and continuously refines service provider scope and performance expectations to meet evolving business needs.

Requirements

Required Education: Bachelor’s Degree

Preferred Education: Master’s Degree

Required Experience: 10+ years Managing large and complex IT Resiliency programs, 24/7 Global Command Center operations, and major incident management operations. and Global Command Center operations, and major incident management operations.

Required Skills, Knowledge and Abilities: Knowledge of industry best practices around IT Resiliency (Disaster Recovery), Global Command Center operations, and Major Incident Management operations using ServiceNow. Expert level knowledge of all aspects of IT Resiliency, Global Command Center Operations, and Major Incident Management systems and procedures, including Application Impact Analysis, Disaster Recovery Tiering and testing cycles, end-to-end tabletop DR exercises, system and application monitoring, on-call scheduling, major incident management-from initation to closure, as it pertains to large and complex university environment, or equivalent. Expert level knowledge of monitoring architecture, solutions, and best practices. Excellent problem-solving, organizational, and communication skills.

Preferred Skills, Knowledge and Abilities: Experience directing the development and execution of tabletop exercises modeling approach IT Disaster scenarios in order to stress test IT DR communications and processes

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on sciencejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:53 min

Applying software development methodologies to incident response

Tobias Dunn-Krahn · LIVE

1:53 min

Managing infrastructure limitations with managed Amazon Aurora databases

Dharin Shah Dharin Shah · World Congress 2025

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

56 sec

Integrating automated approval workflows into the portal

Markus Eisele Markus Eisele · World Congress 2025

5:00 min

Managing complex state with scope-based resource management

Bjarne Stroustrup · World Congress 2022

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · World Congress 2024

Videos

See all

Related articles

See all