> Markdown version of [/jobs/ext/3000172-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3000172-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Fingerprints - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $152,000.0 - $205,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Cloud Computing, Code Coverage, Code Review, Databases, Distributed Systems, Fault Tolerance, Python (Programming Language), Load Testing, Redis, Reliability Engineering, Prometheus, Software Construction, Datadog, Load Balancing, Amazon ElastiCache, Grafana, Containerization, Kubernetes, Terraform, Software Version Control - **Published:** September 19, 2026 - **Apply:** https://www.builtincolorado.com/job/senior-site-reliability-engineer/11262845?handler=ApplyRedirect ## About the Role * 6-10 years of experience in SRE, production engineering, infrastructure, or backend engineering within primarily cloud-based environments (AWS preferred), with meaningful time spent responsible for systems in production. * A track record of owning a system end to end - you've designed something significant, shipped it, operated it, and lived with the consequences when it misbehaved. * Hands-on experience defining and operating against SLIs, SLOs, and error budgets - not just reading the book, but getting targets adopted and acted on. * Strong incident skills: you've led or been a primary responder on high-severity, customer-facing incidents, and you've improved how an organization learns from them. * Depth in distributed systems failure modes in high-throughput, low-latency environments - cache and database saturation, cascading failure, retry storms, capacity limits, degradation and load shedding. * Depth in cloud infrastructure fundamentals: networking, load balancing, containerization (EKS/Kubernetes), and distributed systems. * Strong hands-on experience managing infrastructure through code and configuration (Terraform or equivalent). * Solid programming skills in Go, Python, or a comparable language - you write real, production-ready software and can ship the fix rather than only recommend it. * Fluency with observability tooling (Datadog, Prometheus, Grafana, OpenTelemetry, or similar), including instrumenting systems yourself rather than inheriting dashboards. * Hands-on experience operating Redis/ElastiCache in production - including cluster/shard management, failover behavior, memory eviction policies, and scaling strategies. This is a current skill gap on our team, so depth here is a strong differentiator. * Fluency with software engineering best practices: source control, code review, comprehensive test coverage across edge cases and errors, and safe deployment. * A high level of personal ownership and autonomy, with real experience working without clearly defined requirements. * Pragmatism over purity - you know reliability competes with delivery, can make the case for the right investment at the right time, and can say when a risk is acceptable. * Strong written and verbal communication in English - clear technical design docs, PR reviews, incident updates, and postmortems that bring engineers outside your team along on a decision. * AI-native by default. You use AI tools as a normal part of how you investigate incidents, analyze telemetry, write runbooks, and build tooling - and you have opinions, from experience, about where they help and where they don't yet. ## Description Are you a systems-minded engineer who is happiest when production tells you something you didn't expect? Do you care less about how a system looks on a diagram than about how it behaves at 3am under load it wasn't designed for? Do you want to own reliability for a platform that answers millions of identification requests a day, where being wrong or being slow is a customer-visible event? If so, we have the perfect opportunity for you. We're looking for a Senior Site Reliability Engineer to join our Infrastructure team and take ownership of how our platform behaves in production. This is a hands-on engineering role, not an oversight one - you'll write code and infrastructure, own systems end to end, and be measured by whether the things you own stay fast, available, and predictable as we grow. You'll work across observability, incident response, capacity and performance, change safety, and the tooling that makes all of it routine. You'll define what "reliable" means for the critical paths you own, instrument them so we know before customers do, and partner with product engineering teams to make their services operable by design rather than by heroics. Responsibilities * Own the reliability of core production systems end to end - you instrument them, set targets for them, operate them, and are accountable for how they behave under real traffic. * Define and maintain SLIs and SLOs for the critical paths you own, wire them into dashboards and alerts, and use error budget burn as the evidence base for what gets fixed next. * Drive alert quality: raise signal, kill noise, and close the gap where customers notice a problem before our monitoring does. Anomaly and correctness detection matter as much as uptime. * Take a lead role in incident response - investigate systematically across service boundaries, restore service, and write postmortems that produce follow-ups people actually complete. * Build secure, resilient, and cost-efficient infrastructure, with explicit attention to failure modes: timeouts and retries, backpressure and load shedding, graceful degradation, and blast radius containment. * Do capacity and performance work with real data - load testing, profiling, saturation analysis, and headroom planning ahead of growth rather than after an incident. * Improve change safety: progressive delivery, automated rollback, meaningful pre-production signal, and deployment practices that make shipping boring. * Manage infrastructure through code and configuration (we primarily use Terraform), consistently applying patterns that align with our overall service architecture. * Design, write, and ship software and developer-facing tooling that reduces toil and makes operating services straightforward for the engineers who own them. * Run deliberate failure testing - game days and chaos exercises, staging first - to find the gaps and safe limits before customers do. * Partner with product engineering teams on production readiness for new and high-risk services: capacity, failure modes, rollback plans, runbooks, and on-call handoff. Teach through review rather than gatekeeping. * Participate in the on-call rotation, and improve it: better runbooks, clearer escalation, less pager fatigue for everyone in it. * Approach all engineering work with a security lens - actively looking for vulnerabilities in your own work and in peer reviews. * Act as the go-to person for hard production problems in your area, and mentor engineers through code review, pairing, and design feedback so operational knowledge doesn't silo., For US-based employees, the cash compensation range for this role is $152,000 - $205,000. We set standard ranges for all US roles based on function, level, and geographic location, benchmarked against similar stage growth companies. To comply with local legislation and provide greater transparency, we share salary ranges on all job postings. However, these ranges are specific to the hiring location and may differ within or outside the US. Offers vary depending on, but not limited to, relevant experience, education, certifications/licenses, skills, training, and market conditions. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Top Characteristics of a Software Engineer](https://www.wearedevelopers.com/magazine/166-top-characteristics-of-a-software-engineer) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)