> Markdown version of [/jobs/ext/1459631-major-incident-manager-incident-management-tiktok-usds](https://www.wearedevelopers.com/jobs/ext/1459631-major-incident-manager-incident-management-tiktok-usds). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Major Incident Manager, Incident Management -TikTok USDS - **Company:** Tiktok Usds - **Location:** Seattle, WA, United States - **Experience:** Experienced - **Salary:** $129,960.0 - $246,240.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, JIRA, Microsoft Azure, Bash Shell, Software as a Service, Cloud Computing, Continuous Integration, Data Security, Data Visualization, Python (Programming Language), Ansible, Prometheus, Runbook, SQL Databases, Tableau (Software), Grafana, Mttr, Build Management, Kubernetes, Information Technology, Low-code, Integration Frameworks, Splunk, New Relic (SaaS), Pagerduty, Microservices - **Published:** July 27, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=9301e869ccb9626d ## About the Role Minimum Qualifications - Bachelor's degree in Computer Science, Information Technology, or a related technical field (or equivalent practical experience) with 4+ years of experience in Incident Management, Production Support, or SRE/Operations within a large-scale SaaS or cloud environment. - Strong understanding of cloud computing concepts (AWS, GCP, or Azure) and modern infrastructure architectures (microservices, Kubernetes, CI/CD pipelines). - Hands-on experience configuring and refining alerts and monitoring templates using industry-standard tools (e.g., Grafana, Splunk, Prometheus, New Relic). - Ability to write scripts (e.g., Bash, Python) or use low-code integration tools to automate repetitive incident management tasks and notification pipelines. - Proven ability to remain calm under pressure, command technical bridges, and make decisive, high-stakes calls during complex outages. - Exceptional verbal and written communication skills, with a track record of translating deeply technical issues into clear business impact summaries for leadership. - Strong working knowledge of ITIL incident management frameworks or modern DevOps/SRE incident response practices. Preferred Qualifications - 5+ years of experience in Incident Management, Production Support, or SRE/Operations within a large-scale SaaS or cloud environment. - Deep alignment with Google SRE principles, including Error Budgets, SLA/SLO/SLI management, and automation-first mindsets. - Proven experience implementing ChatOps, auto-remediation workflows, or Event-Driven Ansible/Runbook automation to programmatically resolve common alerts. - Advanced proficiency with modern incident orchestration platforms like PagerDuty, Opsgenie, JIRA Service Management, and Slack integrations. - Experience using SQL, Tableau, or similar data visualization tools to build operational dashboards and report on organizational reliability metrics. - Relevant industry certifications such as ITIL v4, AWS/GCP Certified Professional, or Certified Incident Commander. - Experience coaching and training engineering teams on best practices for on-call hygiene and conducting constructive post-mortems. ## Description The USDS JV Incident Management Team (IMT) is a critical pillar within the US Tech and Product organization, dedicated to ensuring the resilience and reliability of TikTok's U.S. Data Security infrastructure. As a security-first division, we focus on providing specialized oversight and protection for U.S. user data and the platforms that support them. While global teams monitor overall service health, the IMT is uniquely positioned to manage the "blast radius" within the USDS environment. We act as the bridge between technical engineering teams (SRE, Infrastructure, Platform) and business stakeholders, ensuring that every major incident is handled with the precision and urgency required by our unique operating environment. As a Major Incident Manager, you will be at the forefront of protecting the TikTok experience for millions of users. You will join a cohesive, proficient team committed to upholding the highest standards of professionalism and technical expertise, directly contributing to the safety and security of the U.S. digital ecosystem. Responsibilities - Incident Command: Lead the end-to-end response for high-severity incidents, driving rapid restoration of service while coordinating across cross-functional engineering teams. - Stakeholder Communication: Craft and deliver clear, timely, and accurate communication updates to internal stakeholders, executive leadership, and customer-facing teams during critical events. - Post-Incident Reviews: Facilitate blameless post-mortems, ensuring root causes are thoroughly identified, actionable remediation items are tracked, and lessons learned are shared globally. - Dashboarding & Visibility: Create and maintain real-time operational dashboards that provide high visibility into system health, incident trends, and key reliability metrics (MTTD/MTTR). - Continuous Improvement: Analyze incident data and operational metrics to identify systemic trends, driving initiatives to reduce Mean Time to Detect (MTTD) and Mean Time to Resolution (MTTR). - Team Process Evolution: Manage and optimize the incident management rotation, refining playbooks, alerting thresholds, and escalation pathways for the SRE org. - Alerting & Monitoring Improvements: Partner with SRE and development teams to continuously refine alert thresholds, reduce alert fatigue, and ensure high-severity pages are highly actionable. - Automation Engineering: Design and build automated workflows to streamline incident response, such as automated stakeholder communications, auto-remediation scripts, and tool integrations. - Oncall Support: Participate in on-call rotation, including weekends and holidays as needed. Flexibility to join support evening or late night shift. ## Related Videos - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023)