> Markdown version of [/jobs/ext/2416254-principal-engineer-major-incident-response-itil-platform-lead](https://www.wearedevelopers.com/jobs/ext/2416254-principal-engineer-major-incident-response-itil-platform-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Engineer, Major Incident Response & ITIL Platform Lead - **Company:** BJ's Wholesale Club, Inc. - **Location:** Marlborough, MA, United States - **Experience:** Expert - **Salary:** $137,500.0 - $175,500.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), JIRA, Build Automation, Configuration Management Databases, Information Systems, Monitoring of Systems, Python (Programming Language), Routing, Windows PowerShell, Software Engineering, Workflow Management Systems, Working Model 2D, Data Logging, Scripting, Mttr, Build Management, Information Technology, Integration Frameworks, Splunk, Dynatrace, Pagerduty, Servicenow - **Published:** August 24, 2026 - **Apply:** https://www.dice.com/job-detail/662feaaa-1dc2-42fd-8859-c3133c9613e9 ## About the Role Education & Certifications * Bachelor's degree in Computer Science, Information Systems, or equivalent experience. * ITIL 4 Strategic Leader or Managing Professional certification strongly preferred; ITIL Expert acceptable. * ServiceNow certifications (CSA, CIS-ITSM, or CIS-Event Management) highly desirable., * 8+ years in IT Service Management with a strong technical bias - hands-on platform work, not just process governance. * 5+ years of direct experience administering or engineering ServiceNow (workflow design, scripting, integrations, Performance Analytics). * Proven track record leading Major Incident Response in large-scale retail, e-commerce, or distributed digital environments. * Experience owning Problem Management and PIR programs with measurable outcomes (repeat incident reduction, MTTR improvement). * Demonstrated ability to manage and develop onshore/offshore teams in a follow-the-sun operations model. Work Environment: * Hybrid working model: 3 days onsite (Tue, Wed, Thu), 2 days remote (Mon, Fri). * Occasional travel to company locations or industry events. * Flexibility in working hours to accommodate global operations and time zone differences. * Participation in Major Incident on-call rotation. Technical Skills * Deep ServiceNow platform expertise understanding platform mechanics that drive configuration and administration. * Working knowledge of tools such as Dynatrace, Splunk, AlertOps/PagerDuty, or equivalent; ability to build event-to-incident automation bridges. Observability & Monitoring: * MS Teams, Jira - including integration design with ServiceNow. Collaboration & Comms: * ServiceNow Performance Analytics, dashboard design, SLA/SLO instrumentation. Data & Reporting: * Comfort with scripting (JavaScript, Python, or PowerShell) to accelerate toil elimination. Scripting / Automation: Leadership & Soft Skills * Able to command a major incident bridge - calm, decisive, technically credible under pressure. * Executive-level communication: concise, audience-aware, and trustworthy during crises. * Program management discipline: roadmaps, metrics, stakeholder alignment, and backlog ownership. * Growth mindset with a bias toward automation and measurable improvement. ## Description The Principal Engineer, Major Incident Response & ITIL Platform Lead is a senior individual contributor and program leader who combines deep technical expertise with operational discipline. This role owns the design, configuration, and continuous evolution of the ITIL practice - including Major Incident Response (MIR), Post-Incident Review (PIR), and Problem Management - while serving as a hands-on engineer within ServiceNow and adjacent tooling platforms. Unlike a traditional SDM role, this position is explicitly technical: you will architect workflows, build automation, instrument observability, and drive platform maturity across stores, distribution centers, and digital environments. You will also lead a high-performing offshore team and act as the primary program authority during high-severity events - bridging the gap between engineering execution and executive communication., Major Incident Response (MIR) Program Leadership * Own and operate the MIR program end-to-end - from playbook authorship to real-time bridge command - for incidents impacting stores, DCs, POS, fuel, e-commerce, and membership systems. * Serve as Incident Commander during P1/P2 events, driving technical triage, stakeholder communication, and escalation decisions under pressure. * Design and maintain a universal MIR playbook with consistent execution standards 24x7, including on-call rotations for nights, weekends, and holidays. * Establish leadership notification templates, technical bridge protocols, and business-facing communication cadences during major incidents. * Instrument incident severity classification logic, auto-routing, and escalation thresholds directly within ServiceNow. Post-Incident Review & Postmortem Excellence * Own the end-to-end PIR lifecycle - blameless, data-driven reviews completed within SLA - and enforce action-item closure rigor. * Build and maintain an enterprise-wide RCA library, problem signatures, and trend intelligence within ServiceNow's CMDB and Problem Management modules. * Partner with SRE and Software Engineering to translate RCA findings into reliability-driven design improvements and automated runbooks. * Configure and manage PIR workflows, SLA timers, and notification rules natively in ServiceNow - no manual handoffs. ITIL Platform Engineering & ServiceNow Ownership * Act as a hands-on technical owner of ServiceNow ITSM modules: Incident, Problem, Change, and Event Management. * Design and build ServiceNow workflows, business rules, UI policies, Flow Designer automations, and integration spokes connecting monitoring platforms (Dynatrace, Splunk, PagerDuty/AlertOps, etc.). * Develop and maintain custom dashboards, real-time KPI reporting, and SLA/SLO tracking within ServiceNow Performance Analytics. Problem Management * Own the Problem Management lifecycle: identification, logging, root cause investigation, routing, and verified resolution. * Surface recurring incident patterns from trend analysis and feed intelligence back into MIR and Service Excellence programs. * Ensure complete, accurate, and timely documentation of all Problems in ServiceNow with appropriate categorization and linkage to incidents and changes. Continuous Improvement & Team Leadership * Lead and develop a high-performing offshore operations team, setting clear goals aligned to MIR and ITIL program objectives. * Drive a culture of automation-first thinking: identify manual toil and eliminate it through ServiceNow scripting, Flow Designer, and third-party integrations. * Conduct regular retrospectives, process audits, and tooling reviews; translate findings into prioritized improvement backlog items. * Present program health, metrics, and roadmap updates to senior IT and business leadership. Key Outcomes: * Faster stabilization of high-severity events through structured, technically informed incident command. * Measurable reduction in repeat incidents via high-quality, action-tracked RCAs. * A mature, automated ServiceNow platform that minimizes manual effort and accelerates response and reporting. * Predictable, trust-building communication to business stakeholders during and after major incidents. * Continuous improvement embedded into operational DNA - not a periodic exercise. KPIs & Success Metrics: Implement and manage service level agreements (SLAs and SLOs) to meet organizational goals and user expectations. * Mean Time to Acknowledge (MTTA) and Mean Time to Resolve (MTTR) for P1/P2 incidents. * Postmortem SLA compliance (e.g., 100% PIR completion within 5 business days). * Action item closure rate from PIRs within agreed timelines. * Reduction in repeat incidents (measured quarterly). * Problem Management throughput (number of problems logged, analyzed, and resolved). * Leadership communication SLA adherence during major incidents. * Continuous improvement initiatives delivered (e.g., automation, process optimization). * Stakeholder satisfaction scores from incident and problem management processes. ## Related Videos - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023) - [Promotion Interview Questions: How to Answer and Get the Job](https://www.wearedevelopers.com/magazine/413-promotion-interview-questions-how-to-answer-and-get-the-job) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?](https://www.wearedevelopers.com/magazine/701-what-makes-wearedevelopers-world-congress-different-from-every-other-tech-event) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)