> Markdown version of [/jobs/ext/2712994-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2712994-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** RIDGELINERS, LLC - **Location:** Reno, NV, United States - **Experience:** Experienced - **Salary:** $182,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Build Automation, Bash Shell, Mobile Application Development, Software as a Service, Information Systems, Databases, Cursor (Graphical User Interface Elements), DevOps, Github, Identity and Access Management, Python (Programming Language), Node.Js, Reliability Engineering, TypeScript, Circleci, Data Logging, Scripting, Cloud Platform System, Kotlin, Amazon Relational Database Service, Kubernetes, Information Technology, Deployment Automation, Build Tools, Cloudwatch, Terraform, Dynatrace, Golang - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-software-engineer-site-reliability-engineering-ridgeline-8689695 ## About the Role * 8 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or a related discipline. * At least 2 years supporting mission-critical production SaaS workloads running on AWS. * Experience operating production systems where uptime, performance, and reliability are business critical. * Hands-on experience with AWS services including EC2, ECS or EKS, RDS, S3, IAM, CloudWatch, and managed database or messaging services. * Strong understanding of observability, including monitoring, alerting, distributed tracing, and production diagnostics. * Experience designing or significantly improving CI/CD pipelines using tools such as GitHub Actions, CircleCI, Buildkite, or similar platforms. * Experience with deployment strategies including blue/green, canary, or progressive rollouts. * Proficiency in Python, Go, Bash, or another scripting language used for automation and tooling. * Experience implementing Infrastructure as Code using Terraform. * Comfortable participating in an on-call rotation and leading incident response with composure. * Excellent communication skills with the ability to explain technical concepts to both technical and non-technical stakeholders. * Demonstrated ability to make measurable improvements to platform reliability, operational efficiency, or developer productivity. * Strong analytical and troubleshooting skills with a passion for solving complex technical challenges. * A collaborative mindset with a desire to learn, mentor others, and contribute to a positive engineering culture. Bonus * Experience with Kubernetes and Helm. * Familiarity with chaos engineering or fault injection practices. * Experience building or contributing to SLO and error budget programs. * Working knowledge of Kotlin, Node.js, or TypeScript. * Experience supporting highly distributed cloud-native applications. * Bachelor's degree in Computer Science, Information Systems, or a related technical discipline. ## Description Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating complex operational challenges, improving observability, and eliminating manual toil through thoughtful engineering? Are you excited by the opportunity to support mission-critical production systems while collaborating with talented engineers in a fast-moving, innovative environment? If so, we invite you to be a part of our innovative team. As a Site Reliability Engineer, you'll help ensure the reliability, scalability, and operational excellence of Ridgeline's mission-critical SaaS platform. You'll partner closely with product and platform engineers to improve service reliability, accelerate engineering velocity through automation, and build systems that are easier to operate from day one. Our team of engineers are building with cutting-edge technologies-like Claude Code and Cursor-in a fast-moving, creative, progressive work environment. You'll play a key role in advancing our observability, release engineering, incident response, and automation capabilities while contributing measurable improvements to platform stability and developer productivity. At Ridgeline, how we work matters as much as what we build. Ridgeliners act like owners, choose growth over comfort, and communicate with transparency. We assume positive intent, bias toward action, and bring solutions-not just problems. We celebrate wins, learn from setbacks, and thrive in a resilient, collaborative, high-performing culture. If this excites you, we'd love to meet you!, * Improve the reliability, availability, and performance of Ridgeline's mission-critical production SaaS platform. * Build automation that measurably increases engineering velocity while reducing operational toil. * Own and improve production observability through metrics, structured logging, distributed tracing, dashboards, and actionable alerting. * Design and enhance CI/CD pipelines, deployment automation, progressive delivery strategies, and rollback mechanisms. * Define and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budget practices to proactively manage reliability. * Identify capacity constraints and reliability risks before they impact customers. * Participate in an on-call rotation, triaging production issues, coordinating incident response, and driving issues to resolution with very infrequent after-hours support. * Lead blameless postmortems and implement long-term improvements that strengthen platform resilience. * Partner with software engineers on infrastructure design reviews to build highly operable, scalable services. * Develop Infrastructure as Code solutions using Terraform and AWS best practices. * Collaborate across a distributed engineering organization while fostering a culture of ownership, transparency, learning, and continuous improvement. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Kotlin Multiplatform - True power of native code reuse](https://www.wearedevelopers.com/videos/4-kotlin-multiplatform-true-power-of-native-code-reuse) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)