> Markdown version of [/jobs/ext/2217550-software-engineer-site-reliability](https://www.wearedevelopers.com/jobs/ext/2217550-software-engineer-site-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Site Reliability - **Company:** Upstart - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $142,000.0 - $196,600.0 - **Contract:** Temporary contract - **Skills:** JavaScript (Programming Language), Artificial Intelligence, Amazon Web Services, Cloud Computing, Cloud Engineering, Code Generation, Programming Tools, Disaster Recovery, Distributed Systems, Python (Programming Language), Operational Databases, Reliability Engineering, Software Engineering, TypeScript, Datadog, System Availability, Kubernetes, Infrastructure Automation Frameworks, Sumo Logic (Software), Cloudwatch, Golang - **Published:** August 25, 2026 - **Apply:** https://www.dice.com/job-detail/067ea175-9c6e-460f-9d7b-7e67aa404917 ## About the Role * 3+ years of professional experience in software engineering, site reliability engineering, or a related engineering discipline * Strong software development skills in one or more general purpose programming languages such as Python, Go, JavaScript, or TypeScript * Experience designing, building, testing, and operating production software, internal tooling, or infrastructure * Experience with cloud infrastructure, distributed systems, observability, monitoring, or production operations * Experience participating in on call or incident response for production systems * Demonstrated ability to independently deliver well scoped engineering projects, navigate technical ambiguity, and collaborate effectively across engineering teams * Demonstrated experience using AI assisted development tools across multiple stages of the software engineering lifecycle, with an interest in continually evolving how you use these tools as their capabilities advance Preferred Qualifications * Experience with Kubernetes, AWS, infrastructure as code, and cloud native production environments * Experience building internal reliability, observability, incident management, or operational automation tools * Experience with observability platforms such as Datadog, Sumo Logic, CloudWatch, or similar technologies * Experience with reliability practices such as service level objectives, capacity planning, resiliency testing, disaster recovery, or operational readiness * Experience operating distributed applications with complex dependencies and high availability requirements * Experience building AI enabled operational workflows, tools, or automation that extend beyond individual code generation ## Description As a Software Engineer on the Site Reliability Engineering team, you will build and operate systems that improve the reliability, resiliency, and observability of Upstart's production environment. You will work across observability platforms, reliability tooling, incident response systems, operational automation, and resiliency capabilities. You will independently deliver well scoped engineering projects, contribute to technical design, and use production data and operational experience to improve systems used across engineering. We are looking for engineers who are thoughtful and intentional about how AI changes software development and operations. You should be comfortable using AI throughout the engineering lifecycle, including understanding unfamiliar systems, investigating production behavior, planning implementation, accelerating development, validating changes, and automating repetitive work. We expect engineers to continually develop more effective ways of working with increasingly capable AI tools and to apply sound engineering judgment to where they provide the most leverage. How you'll make an impact * Build and improve the tooling, services, and automation that help engineers understand and improve the reliability of production systems * Develop shared observability capabilities that make metrics, logs, traces, service health, and customer impact easier to understand and act on * Improve incident response and operational readiness through better tooling, automation, standards, and actionable production signals * Build resiliency capabilities that help teams identify failure modes, reduce operational risk, and recover effectively from infrastructure or application failures * Identify recurring operational toil and reliability problems and replace manual processes with durable software and automation * Use AI as an integrated part of software development and operational problem solving, while identifying opportunities for AI enabled capabilities that improve incident investigation, observability, reliability, and engineering efficiency ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From Black Box to Glass Box : Bedrock AgentCore Observability](https://www.wearedevelopers.com/videos/2126-from-black-box-to-glass-box-bedrock-agentcore-observability) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)