> Markdown version of [/jobs/ext/533912-senior-site-reliability-engineer-platform-responsibility-usds](https://www.wearedevelopers.com/jobs/ext/533912-senior-site-reliability-engineer-platform-responsibility-usds). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer, Platform Responsibility - USDS - **Company:** Tiktok Inc. - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $177,688.0 - $341,734.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Big Data, Unix, Software Quality, DevOps, Disaster Recovery, Distributed Systems, Log Analysis, Networking Basics, Reliability Engineering, Prometheus, Datadog, Large Language Models, Grafana, Kubernetes, Information Technology, Apache Kafka - **Published:** June 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8f0513ed48e549ff ## About the Role Do you have experience in UNIX?, Do you have a Bachelor's degree?, Minimum Qualifications: - Bachelor's degree or above in Computer Science or a related technical discipline. - At least 5 years of professional experience in SRE, DevOps, or Infrastructure Engineering. - Proven experience integrating AI/LLM APIs into internal workflows, specifically for log analysis, alert contextualization, or diagnostic assistance. - Deep understanding of Unix/Linux system internals, networking fundamentals, and distributed systems architecture. - Expertise in designing and scaling observability stacks using tools such as Prometheus, Grafana, or DataDog. - Demonstrated ability to troubleshoot complex, non-obvious production issues across the entire stack. Preferred Qualifications: - Experience building agentic workflows and orchestration frameworks to assist in incident triaging and runbook matching. - Mastery of AI-assisted development tools to accelerate infrastructure-as-code delivery and automate documentation. - Deep technical proficiency in container orchestration via Kubernetes and managing big data technologies such as Kafka. ## Description administration and operational efficiency, leveraging AI-assisted development tools to accelerate delivery and code quality - Participate in regular on-call rotations as part of a team that provides 24 hour coverage across multiple shifts - Engage in and improve the whole lifecycle of services from inception and design, development, capacity planning, and launch reviews, to deployment, operation, and refinement - Practice sustainable user support, incident response, and post mortem - Drive the long-term roadmap for system administration tools. Build internal platforms that leverage AI-assisted development to eliminate toil and improve engineering velocity. - Serve as a primary Incident Commander for high-severity issues, leading cross-functional teams and ensuring technical resolution aligns with business priorities. - Partner with Product and Development teams from the design phase to ensure observability, scalability, and disaster recovery are core components of every new feature. - Foster a culture of sustainable operations by mentoring junior engineers and evangelizing SRE principles throughout the organization. ## Related Videos - [WeAreDevelopers LIVE - Node and Package Security](https://www.wearedevelopers.com/videos/2138-wearedevelopers-live-node-and-package-security) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)