> Markdown version of [/jobs/ext/2677516-director-major-incident-management](https://www.wearedevelopers.com/jobs/ext/2677516-director-major-incident-management). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Director, Major Incident Management - **Company:** Visa Inc. - **Location:** Denver, CO, United States (Remote available) - **Experience:** Expert - **Salary:** $160,100.0 - $256,300.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cyber Security, Distributed Systems, Monitoring of Systems, Reliability Engineering, Generative AI - **Published:** September 2, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/18163939?backUrl=%2Fcareer%2F18163939%2FDirector-Major-Incident-Management-Colorado-Denver ## About the Role * 10+ years of relevant work experience with a Bachelor's Degree or at least 7 years of work experience with an Advanced degree (e.g. Masters, MBA, JD, MD) or 4 years of work experience with a PhD, OR 13+ years of relevant work experience., * 12 or more years of work experience with a Bachelor's Degree or 8-10 years of experience with an Advanced Degree (e.g. Masters, MBA, JD, MD) or 6+ years of work experience with a PhD * 10+ years of relevant work experience with a Bachelor's Degree, OR 13+ years of relevant work experience. * 12+ years of experience in technology operations, infrastructure operations, site reliability engineering, incident management, service management, or related disciplines. * 5+ years of people leadership experience managing high-performing technical teams. * Proven experience leading the response and recovery of large-scale production incidents. * Experience working in complex, mission-critical, global technology environments. * Deep knowledge of Major Incident Management, IT Service Management (ITSM), Change Management, Problem Management, and Operational Governance. * Experience supporting highly available distributed systems with stringent service availability requirements. * Strong understanding of cloud platforms, observability solutions, monitoring technologies, and modern operational practices. * Experience with SRE principles, operational resilience frameworks, AIOps, and automation technologies. * Demonstrated success influencing senior executives and cross-functional technology leaders. * Exceptional written, verbal, and executive communication skills. * Experience operating in regulated or highly secure environments. * Knowledge of AI, Generative AI, automation, or predictive operations capabilities is strongly preferred. ## Description Visa is seeking an experienced and highly effective technology operations leader to serve as Director, Major Incident Management, with enterprise accountability for Visa's global major incident response strategy, operating model and execution. This role is responsible for leading Visa's global Major Incident Management function, driving the rapid mitigation and resolution of critical technology incidents impacting Visa's products, services, infrastructure, and clients. The successful candidate will lead a globally dispersed team responsible for influencing cross-functional incident response across engineering, infrastructure, security, product, and operations organizations while maintaining executive-level visibility and stakeholder confidence during high-severity events. This leader will define and advance operational resilience, incident governance, automation, and AI-enabled operational capabilities within one of the world's most complex and highly available technology environments. Responsibilities include executive communications, incident command, operational readiness, post-incident review governance, and continuous improvement of incident response processes. Essential Functions Major Incident Leadership * Own Visa's global Major Incident Management organization supporting a 24x7x365 technology environment. * Lead enterprise-wide escalation, and management of business-critical technology incidents with significant client, operational, reputational and regulatory impact. * Establish operational priorities and drive rapid service restoration during high-severity incidents. * Serve as the senior accountable escalation point for complex and cross-functional production events. * Ensure incident response activities maintain an appropriate balance between speed, risk management, and business impact mitigation. Executive Communications and Stakeholder Engagement * Lead communication strategies during major incidents, ensuring timely, accurate, and concise updates for executive leadership, technology partners, and key stakeholders. * Provide executive summaries, operational risk assessments, and incident status communications. * Build trusted relationships across Product Development, Engineering, Infrastructure, Security, Operations, and Corporate Functions. Operational Excellence and Governance * Define, implement, govern and continuously improve incident management standards, operating procedures, policies, and performance metrics. * Establish and monitor key performance indicators supporting operational effectiveness and service resilience. * Lead operational reviews, trend analysis, and improvement initiatives designed to reduce customer impact and operational risk. * Drive operational discipline and accountability across incident response stakeholders. Post-Incident Review and Continuous Improvement * Establish governance for root cause analysis and corrective action programs. * Ensure lessons learned are translated into measurable operational improvements. * Partner with engineering and operations leaders to address recurring issues and systemic weaknesses. * Champion a culture of learning, accountability, and continuous improvement. Operational Resilience and Readiness * Partner with crisis management, business continuity, cybersecurity, and operational resilience functions. * Design and lead enterprise major incident simulations, readiness exercises, and operational reviews. * Ensure organizational preparedness for large-scale technology disruptions and high-impact events. Automation and AI-Enabled Operations * Drive the adoption of automation, advanced analytics, and AI-driven capabilities across incident management processes. * Partner with engineering and automation teams to improve incident detection, triage, correlation, response, and communication workflows. * Identify opportunities to reduce manual effort while improving operational speed, consistency, and quality. * Drive the evolution toward predictive and proactive operational models. People Leadership * Build, develop, and lead a diverse, high-performing global team. * Foster a culture centered on operational excellence, collaboration, resilience, accountability, and innovation. * Coach and mentor future leaders while supporting career development and succession planning. Visa requires at least 3 days in office, expectations of these days will be confirmed by your Hiring Manager., * In this role, you will: * Improve service restoration speed and operational responsiveness. * Reduce repeat incidents and customer-impacting disruptions. * Increase automation and operational efficiency. * Strengthen executive confidence during critical events. * Enhance operational resilience across Visa's technology ecosystem. * Build a world-class incident management organization recognized as a strategic partner to engineering and product teams. ## Related Videos - [Your imaginations is (no longer) the limit: how Generative AI empowers people to be creative](https://www.wearedevelopers.com/videos/741-your-imaginations-is-no-longer-the-limit-how-generative-ai-empowers-people-to-be-creative) - [Thinking Differently - How to Make Money from Cyber Attacks & Cheats](https://www.wearedevelopers.com/videos/745-thinking-differently-how-to-make-money-from-cyber-attacks-cheats) - [Best Practices for AI-Assisted Development of Distributed Systems](https://www.wearedevelopers.com/videos/100200-best-practices-for-ai-assisted-development-of-distributed-systems) - [Your next 10x engineer isn't in your city. Refactor accordingly.](https://www.wearedevelopers.com/videos/100103-your-next-10x-engineer-isn-t-in-your-city-refactor-accordingly) - [The shadows that follow the AI generative models](https://www.wearedevelopers.com/videos/624-the-shadows-that-follow-the-ai-generative-models) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)