> Markdown version of [/jobs/ext/1347551-senior-engineering-manager-site-reliability](https://www.wearedevelopers.com/jobs/ext/1347551-senior-engineering-manager-site-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Engineering Manager, Site Reliability - **Company:** Upstart - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $195,300.0 - $270,400.0 - **Contract:** Temporary contract - **Skills:** Amazon Web Services, Cloud Computing, Distributed Systems, Operational Data Store, Reliability Engineering, Prometheus, Software Engineering, Datadog, Grafana, Kubernetes - **Published:** July 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a6b1a70eac4a83ea ## About the Role If you're energized by tackling meaningful problems, excited to innovate with purpose, and motivated by work that truly matters, we'd love to hear from you., * 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering * Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems * Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes * Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations * Experience leading high severity incident response and improving incident management practices at scale * Demonstrated ability to translate strategy into focused, capacity aware plans and deliver measurable outcomes * Track record of hiring, developing, and retaining high performing engineers and engineering leaders * Strong cross-functional leadership and communication, with the ability to turn complex operational data into clear decisions and drive alignment across teams Preferred Qualifications * Experience operating large scale, highly available distributed systems * Experience implementing or evolving service-level objectives and error-budget practices * Experience with observability platforms such as Datadog, Grafana, Prometheus, OpenTelemetry, or similar technologies * Experience developing incident management, operational readiness, or resilience programs across a large engineering organization * Familiarity with Kubernetes, AWS, and modern cloud native architectures * Experience supporting major platform or architectural transitions * Strong product mindset when building internal reliability capabilities * Experience establishing executive level reliability reporting and operating reviews ## Description The Site Reliability Engineering (SRE) team enables Upstart's engineering organization to operate reliable, observable, and resilient systems at scale. The team owns company-wide incident management practices, reliability standards, operational readiness, and the capabilities that help engineering teams identify, respond to, and learn from production issues. Our goal is to make reliability an integrated part of how software is designed, delivered, and operated. We are building a model where engineering teams have the trusted signals, automated safeguards, and operational practices needed to move quickly while protecting our customers and business. The team advances observability, incident detection and response, service level objectives, operational readiness, and systemic improvements based on incident learnings. SRE partners across product engineering, infrastructure, security, and platform teams to improve reliability at scale. The Role: As the Senior Engineering Manager of Site Reliability Engineering, you will lead a team responsible for improving the reliability and operational maturity of Upstart's products and services. You will drive high impact improvements across incident management, observability, operational readiness, and reliability engineering. You will serve as the accountable leader of the SRE function, translating reliability strategy into focused plans, clear ownership, and measurable outcomes. You will partner closely with engineering leaders to establish reliability expectations, identify systemic risks, and build scalable capabilities that enable teams to operate services safely and independently. This role is suited for a hands-on leader with strong technical judgment, disciplined execution, and a track record of building high performing teams. You will balance immediate operational needs with durable improvements that reduce risk, strengthen resilience, and improve how Upstart learns from production., * Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering * Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function * Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures * Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts * Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership * Set a high bar for technical quality, operating rigor, and executive communication * Develop engineers and leaders who can independently own complex reliability initiatives, * Evolve Upstart's incident management program to improve detection, response, coordination, communication, and recovery * Establish clear standards for managing high severity incidents and provide visible leadership during critical events * Improve postmortem quality and ensure incident learnings result in durable engineering improvements * Identify recurring failure patterns and drive systemic solutions across teams * Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction, * Improve the quality, accessibility, and trustworthiness of signals used to understand production health * Drive consistent practices across metrics, logs, traces, alerting, and service health * Advance the use of service level objectives and customer impact signals to guide priorities and operational decisions * Reduce detection gaps, noisy alerts, manual investigation, and recurring operational toil * Define measurable reliability outcomes and use data to prioritize investments and communicate impact * Partner with platform and product engineering teams to embed reliability into standard engineering workflows Operational Readiness and Resilience * Establish scalable operational readiness standards for new services, major launches, and architectural changes * Set clear expectations for service ownership, monitoring, capacity, failure handling, and incident response * Identify systemic reliability risks and partner with engineering teams to prioritize and address them * Improve resilience through automation, failure testing, recovery capabilities, and operational safeguards * Build operating mechanisms that turn reviews and analysis into clear decisions, owners, timelines, and sustained follow through * Align stakeholders and dependencies before critical launches and engineering decisions ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Data binning and understanding histograms](https://www.wearedevelopers.com/videos/2086-data-binning-and-understanding-histograms) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too)