> Markdown version of [/jobs/ext/2828864-it-incident-manager](https://www.wearedevelopers.com/jobs/ext/2828864-it-incident-manager). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # IT Incident Manager - **Company:** Everforth Apex - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Distributed Systems, Mainframes, Software Engineering - **Published:** September 10, 2026 - **Apply:** https://www.dice.com/job-detail/d2ea234b-14f3-464d-8af5-13d90c6cc060 ## About the Role * Bachelor's degree in a related business or science field, or equivalent work experience. Work Experience * 3+ years of experience leading technical projects, incidents, or cross-functional initiatives. * Experience managing highly impactful incidents across multiple departments or lines of business. * Experience with distributed systems, networks, application development, and mainframe environments. * Familiarity with ITIL processes and incident management principles. Certifications & Licenses * Certifications in ITIL, incident management, or related disciplines. (Preferred) ## Description The IT Incident Manager leads cross-functional investigative teams to resolve technology events impacting KeyBank's enterprise systems. This individual contributor role is responsible for monitoring and driving incident resolution efforts, proactively engaging technical support partners upon detection of incidents, facilitating communication among technical and business stakeholders, and ensuring timely and effective recovery. The Incident Manager must quickly assess complex technical environments and actively guide troubleshooting efforts. This position includes on-call and off-hours support on a rotating basis, with a standard shift ranging between 11:00 AM-9:00 PM EST. Essential Functions * Identify and assess critical incidents and outages impacting KeyBank's operations. * Proactively and directly investigate potential technology incidents when detected or observed. * Lead incident recovery efforts, coordinating cross-functional teams and vendors. * Engage in independent decision-making and be confident in making judgment calls when necessary. * Facilitate technical crisis calls and ensure appropriate resources are engaged. * Mediate high-stakes and time-sensitive triage bridges and chats to effectively and respectfully drive resolution. * Provide accurate and timely communications to stakeholders and executive leadership. * Escalate critical issues and manage progress updates throughout troubleshooting and remediation. * Lead post-mortem sessions and document incident restoration activities and timelines. * Identify action items and follow-up activities needed following incident resolution and share with problem management. * Ensure recovery documentation is accurate and tested for feasibility. * Drive continuous improvement initiatives to enhance system stability and client experience. * Identify gaps in recovery processes and collaborate with monitoring and detection teams. * Participate in incident management performance reviews and process evaluations. * Ensure adherence to ITSM processes and define/manage critical success metrics. * Support production readiness for critical projects, focusing on incident management. ## Related Videos - [Best Practices for AI-Assisted Development of Distributed Systems](https://www.wearedevelopers.com/videos/100200-best-practices-for-ai-assisted-development-of-distributed-systems) - [The Avengers Initiative (Practical Ethics for Software Engineers)](https://www.wearedevelopers.com/videos/2070-the-avengers-initiative-practical-ethics-for-software-engineers) - [Unlocking the Power of the Mainframe: Developing modern applications on z/OS](https://www.wearedevelopers.com/videos/764-unlocking-the-power-of-the-mainframe-developing-modern-applications-on-z-os) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Your Distributed System Just Got a Brain. Now What?](https://www.wearedevelopers.com/videos/100017-your-distributed-system-just-got-a-brain-now-what) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) ## Related Articles - [The Geometry of Incidents: Connecting User Impact to Architecture](https://www.wearedevelopers.com/magazine/764-the-geometry-of-incidents-connecting-user-impact-to-architecture) - [From developer to manager – what does it take to become an engineering manager?](https://www.wearedevelopers.com/magazine/42-from-developer-to-manager-what-does-it-take-to-become-an-engineering-manager) - [How should you format your IT resume?](https://www.wearedevelopers.com/magazine/68-how-should-you-format-your-it-resume) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)