> Markdown version of [/jobs/ext/2718558-staff-incident-manager](https://www.wearedevelopers.com/jobs/ext/2718558-staff-incident-manager). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Incident Manager - **Company:** YUNO INC. - **Location:** Juno, WA, United States (Remote available) - **Contract:** Permanent contract - **Skills:** Distributed Systems, PCI Data Security Standards, Site Reliability Engineering Practices, Datadog, Data Logging, JavaScript Pagination Plugin, Mttr, Pagerduty - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-incident-manager-yuno-8707489 ## About the Role * Proven experience running major incident response in a production environment that real customers depend on, ideally as an incident commander or in a dedicated incident management function. * Strong working knowledge of modern distributed systems and cloud-native operations - comfortable holding your own on a bridge call with senior engineers during an outage, following the technical thread, and keeping the response moving without needing every detail spoon-fed. * Experience defining or maturing an incident management practice from the ground up: severity frameworks, operating models, runbooks, and on-call programs built to last, not just inherited. * Fluency with observability and incident tooling - metrics, logging, tracing, alerting, and paging platforms. Hands-on experience with Datadog is strongly preferred; familiarity with a paging platform (PagerDuty, OpsGenie, or similar) and a status page tool is expected. * A track record of running postmortems that change behavior, and of closing the loop between incidents and engineering work. * Excellent written and verbal communication - able to write a clear merchant-facing status update and a precise internal escalation under pressure. * Calm, decisive judgment during high-pressure events - you hold structure when others are reacting. * Comfort being on-call as a regular part of the role, including for critical incidents outside business hours. * English - advanced proficiency required., * Familiarity with PCI-DSS and the operational obligations that come with handling payment flows at scale. * Spanish - business-level proficiency preferred; Yuno's engineering organization spans LatAm, and on-call bridges and postmortem discussions often run in Spanish., * Prior experience in payments, fintech, or another high-availability, regulated domain. * Exposure to SRE practices, error budgets, and SLO-driven prioritization. * Experience managing incidents across multi-timezone, remote-first engineering organizations. ## Description We are hiring a Staff Incident Manager to own the major incident lifecycle end to end. You are the person who takes control when something is broken in production, brings the right people together, drives the response to resolution, and makes sure we are measurably better after every incident than we were before it. This is not a passive coordination role. You set the standard for how Yuno detects, responds to, communicates about, and learns from incidents. You will spend as much time fixing the process as you do running the response. On-call responsibility is core to this role. Yuno's engineering teams run a "You Build It, You Run It" model - pods own their own on-call. The Incident Commander function exists to coordinate cross-domain and Sev-1 events where multiple pods are involved., * Incident command - act as incident commander on major and critical incidents, owning coordination, decision-making cadence, and escalation from detection through resolution; treat merchant transaction impact, PSP and acquirer dependency failures, settlement and reconciliation knock-on effects, and PCI-DSS scope as first-class concerns in every response. * Reliability metrics - drive down time to detect, time to engage, and time to recover; own MTTR as a headline metric and the operational discipline behind Yuno's 99.99% uptime target. * On-call program - run rotations, escalation policies, paging hygiene, and alert quality; reduce noise so on-call engineers trust their pages. * Incident communications - own internal stakeholder alignment during an event and drive clear, accurate merchant-facing updates - including the status page - in coordination with Support and account teams. * Postmortems - run blameless postmortems, hold the room to a no-blame standard, and make sure action items are concrete, owned, and tracked to closure. * Reliability roadmap - translate recurring incident patterns into reliability work, partnering with engineering teams to turn postmortem findings into roadmap items, not orphaned tickets. * Operating model - define and maintain incident severity levels, response runbooks, and the operating model for declaring and managing incidents. * Reporting - report on incident trends, reliability posture, and SLA and SLO performance to engineering leadership. ## Related Videos - [AI in Leadership: How Technology is Reshaping Executive Roles](https://www.wearedevelopers.com/videos/1705-ai-in-leadership-how-technology-is-reshaping-executive-roles) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)