Director, Major Incident Management
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Visa is seeking an experienced and highly effective technology operations leader to serve as Director, Major Incident Management, with enterprise accountability for Visa’s global major incident response strategy, operating model and execution.
This role is responsible for leading Visa’s global Major Incident Management function, driving the rapid mitigation and resolution of critical technology incidents impacting Visa’s products, services, infrastructure, and clients. The successful candidate will lead a globally dispersed team responsible for influencing cross-functional incident response across engineering, infrastructure, security, product, and operations organizations while maintaining executive-level visibility and stakeholder confidence during high-severity events.
This leader will define and advance operational resilience, incident governance, automation, and AI-enabled operational capabilities within one of the world’s most complex and highly available technology environments. Responsibilities include executive communications, incident command, operational readiness, post-incident review governance, and continuous improvement of incident response processes.
Essential Functions
Major Incident Leadership
- Own Visa’s global Major Incident Management organization supporting a 24x7x365 technology environment.
- Lead enterprise-wide escalation, and management of business-critical technology incidents with significant client, operational, reputational and regulatory impact.
- Establish operational priorities and drive rapid service restoration during high-severity incidents.
- Serve as the senior accountable escalation point for complex and cross-functional production events.
- Ensure incident response activities maintain an appropriate balance between speed, risk management, and business impact mitigation.
Executive Communications and Stakeholder Engagement
- Lead communication strategies during major incidents, ensuring timely, accurate, and concise updates for executive leadership, technology partners, and key stakeholders.
- Provide executive summaries, operational risk assessments, and incident status communications.
- Build trusted relationships across Product Development, Engineering, Infrastructure, Security, Operations, and Corporate Functions.
Operational Excellence and Governance
- Define, implement, govern and continuously improve incident management standards, operating procedures, policies, and performance metrics.
- Establish and monitor key performance indicators supporting operational effectiveness and service resilience.
- Lead operational reviews, trend analysis, and improvement initiatives designed to reduce customer impact and operational risk.
- Drive operational discipline and accountability across incident response stakeholders.
Post-Incident Review and Continuous Improvement
- Establish governance for root cause analysis and corrective action programs.
- Ensure lessons learned are translated into measurable operational improvements.
- Partner with engineering and operations leaders to address recurring issues and systemic weaknesses.
- Champion a culture of learning, accountability, and continuous improvement.
Operational Resilience and Readiness
- Partner with crisis management, business continuity, cybersecurity, and operational resilience functions.
- Design and lead enterprise major incident simulations, readiness exercises, and operational reviews.
- Ensure organizational preparedness for large-scale technology disruptions and high-impact events.
Automation and AI-Enabled Operations
- Drive the adoption of automation, advanced analytics, and AI-driven capabilities across incident management processes.
- Partner with engineering and automation teams to improve incident detection, triage, correlation, response, and communication workflows.
- Identify opportunities to reduce manual effort while improving operational speed, consistency, and quality.
- Drive the evolution toward predictive and proactive operational models.
People Leadership
- Build, develop, and lead a diverse, high-performing global team.
- Foster a culture centered on operational excellence, collaboration, resilience, accountability, and innovation.
- Coach and mentor future leaders while supporting career development and succession planning.
Visa requires at least 3 days in office, expectations of these days will be confirmed by your Hiring Manager., * In this role, you will:
- Improve service restoration speed and operational responsiveness.
- Reduce repeat incidents and customer-impacting disruptions.
- Increase automation and operational efficiency.
- Strengthen executive confidence during critical events.
- Enhance operational resilience across Visa’s technology ecosystem.
- Build a world-class incident management organization recognized as a strategic partner to engineering and product teams.
Requirements
- 10+ years of relevant work experience with a Bachelor’s Degree or at least 7 years of work experience with an Advanced degree (e.g. Masters, MBA, JD, MD) or 4 years of work experience with a PhD, OR 13+ years of relevant work experience., * 12 or more years of work experience with a Bachelor’s Degree or 8-10 years of experience with an Advanced Degree (e.g. Masters, MBA, JD, MD) or 6+ years of work experience with a PhD
- 10+ years of relevant work experience with a Bachelor’s Degree, OR 13+ years of relevant work experience.
- 12+ years of experience in technology operations, infrastructure operations, site reliability engineering, incident management, service management, or related disciplines.
- 5+ years of people leadership experience managing high-performing technical teams.
- Proven experience leading the response and recovery of large-scale production incidents.
- Experience working in complex, mission-critical, global technology environments.
- Deep knowledge of Major Incident Management, IT Service Management (ITSM), Change Management, Problem Management, and Operational Governance.
- Experience supporting highly available distributed systems with stringent service availability requirements.
- Strong understanding of cloud platforms, observability solutions, monitoring technologies, and modern operational practices.
- Experience with SRE principles, operational resilience frameworks, AIOps, and automation technologies.
- Demonstrated success influencing senior executives and cross-functional technology leaders.
- Exceptional written, verbal, and executive communication skills.
- Experience operating in regulated or highly secure environments.
- Knowledge of AI, Generative AI, automation, or predictive operations capabilities is strongly preferred.
Benefits & conditions
$160,100.00 to $ 256,300.00 life insurance, paid time off, 401(k) United States, Colorado, Denver Remote Work Site (Show on map) Sep 01, 2026, For roles located in the US, the estimated salary range for this position is $160,100.00 to $ 256,300.00 USD per year, which may include potential sales incentive payments (if applicable). Salary may vary depending on job-related factors which may include knowledge, skills, experience, and location. In addition, this position may be eligible for bonus and equity.Visa has a comprehensive benefits package for which this position may be eligible that includes Medical, Dental, Vision, 401(k), FSA/HSA, Life Insurance, Paid Time Off, and Wellness Program.
About the company
Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.
At Visa, you’ll have the opportunity to create impact at scale - tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.
Join Visa and do work that matters - to you, to your community, and to the world. Progress starts with you.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Navigating the AI Shift
How to Become an AI Engineer
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
What Industries Outside of AI Are Hiring The Most AI Experts?