IT Infrastructure & Disaster Recovery Project Manager

Randstad
Houston, TX, United States
7 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$187,200.0 - $197,600.0
Working hours
Regular working hours
Job source

Tech stack

Microsoft Excel Agile Methodology Microsoft Azure Databases Disaster Recovery Executive Information Systems Microsoft Office Microsoft Project Microsoft Visio Team Foundation Server Microsoft SharePoint Virtualization Technology
+2 more
Information Technology Network Server

Job description

An experienced Senior IT Infrastructure and Disaster Recovery Project Manager is needed to support the continued maturation of a large scale enterprise Disaster Recovery program. This role will coordinate production recovery testing, infrastructure and application teams, program governance, reporting, and continuous improvement initiatives designed to strengthen enterprise recovery readiness.

The ideal candidate brings at least seven years of complex IT project management experience, including three to five years supporting infrastructure initiatives and hands-on experience coordinating Disaster Recovery testing, production failover/failback, or comparable complex production cutovers. Candidates should have sufficient technical infrastructure knowledge to effectively coordinate server, database, virtualization, network, application, integration, and operations teams.

The Project Manager will partner closely with Business Continuity and PMO leadership to standardize DR testing processes, improve program governance, track application recovery readiness, capture lessons learned, develop dashboards and metrics, and establish repeatable processes across technical teams.

This is a 100% onsite position in Houston and requires occasional schedule flexibility to support after hours Disaster Recovery testing., * Partner with the Business Continuity Manager and PMO to support and mature the enterprise Disaster Recovery program.

  • Plan, coordinate, and help execute DR tests, including production failover and failback activities across application and infrastructure teams.
  • Standardize DR testing procedures, playbooks, checklists, and governance to create a consistent, repeatable process across PMs and technical teams.
  • Track annual DR testing schedules, application readiness/compliance, risks, issues, dependencies, and outstanding activities.
  • Coordinate across application owners, servers, databases, virtualization, networks, operations, and other infrastructure teams during recovery testing.
  • Capture lessons learned from DR exercises and drive process and program improvements.
  • Develop and maintain program metrics, dashboards, status reporting, and executive/steering committee updates.
  • Help increase application-owner accountability for DR readiness and support greater automation of the recovery/testing process.
  • Support occasional after-hours DR tests while self-managing the work schedule within approximately a 40-hour week.

Requirements

Minimum of 7 years of experience managing multiple, complex IT projects.

3-5 years of experience executing IT infrastructure projects.

Experience supporting Disaster Recovery, Business Continuity, infrastructure resiliency, or Operational Readiness initiatives.

Experience coordinating and executing Disaster Recovery testing involving production failover/failback, production cutovers, or comparable complex production recovery activities.

Strong understanding of IT infrastructure environments and the ability to effectively coordinate server, database, virtualization, network, application, integration, and operations teams.

Demonstrated experience leading cross-functional technical project teams.

Strong governance and process-development experience, including standards, controls, checklists, and repeatable processes.

Strong executive level reporting, dashboard, metrics, and stakeholder management experience.

Advanced Microsoft Excel skills.

Strong project planning, resource management, risk management, dependency management, and problem solving capabilities.

Strong leadership, interpersonal, negotiation, influencing, organizational, and facilitation skills.

Effective written and verbal communication skills.

Experience working with Waterfall and Agile methodologies.

Proficiency with Microsoft Office and project management tools.

Bachelor’s degree in Computer Science, MIS, CIS, a related field, or equivalent professional experience.

Ability to work 100% onsite in Houston.

Ability to occasionally support Disaster Recovery testing outside normal business hours with corresponding flexibility in the regular work schedule.

Preferred Qualifications

Experience supporting a large scale enterprise Disaster Recovery or Business Continuity program.

Experience coordinating complex production failover and failback testing.

Experience developing or improving Disaster Recovery testing processes and recovery playbooks.

Experience with SOX related applications or regulated environments.

Experience with Disaster Recovery automation or recovery-process automation.

Experience developing executive dashboards and program-level reporting.

Experience with SharePoint, MS Project, Azure DevOps/TFS, and Visio.

Project Management Professional (PMP) certification.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:24 min

Estimating project expenditures with Azure Pricing Calculator

Radu Vunvulea Radu Vunvulea · World Congress 2022

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

1:27 min

Binding executable logic into network servers

Saoussen Chaabnia Saoussen Chaabnia · Europe 2026 Virtual

5:31 min

Applying the recovery framework to a management case study

Aleksandra Lemańska Aleksandra Lemańska · Europe 2026 Virtual

Videos

See all

Related articles

See all