> Markdown version of [/jobs/ext/260402-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/260402-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - **Company:** Earnin - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $252,000.0 - $308,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Cursor (Graphical User Interface Elements), Noise Reduction, Distributed Systems, Amazon DynamoDB, Fault Tolerance, Python (Programming Language), Reliability Engineering, Software Engineering, Datadog, Large Language Models, Amazon Relational Database Service, Kubernetes, Apache Kafka, Build Tools, Cloudwatch, Amazon Simple Queue Service (SQS), Terraform - **Published:** May 23, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=fef1827c15f7108d ## About the Role Do you have experience in SAFe?, * 7+ years in SRE, Software Engineering, or Infrastructure Engineering with increasing scope and cross-org influence. Track record of KPI driven reliability and operational excellence improvements at scale. * Demonstrated experience applying AI/LLMs to operational workflows in production: alert triage/resolution, runbook automation, incident investigation, postmortem, or agentic operations tooling. Not a theoretical interest, but shipped work. * Significant expertise with SLOs/SLIs, error budgets, incident command, and blameless postmortems in large-scale distributed systems. You have driven follow-through that actually prevented recurrence. * Meaningful software engineering ability (Python, Go, or similar). You build tools and automation, not just dashboards. * Deep observability experience (Datadog, CloudWatch, OpenTelemetry) with pragmatic, signal-heavy alerting designed for real human response, enhanced by AI-driven noise reduction. * Solid infrastructure-as-code proficiency (Terraform, Kubernetes, AWS) with safe, reversible deployment practices. * Proficiency with AI-assisted development tools (Cursor, Claude Code, Copilot) to accelerate your own engineering work and to model that behavior for the teams you partner with, and experience using AI-assisted development tools as part of your software development workflow * Experience in fintech or regulated environments (SOC 2, PCI), and familiarity with FinOps or cost/performance tradeoffs in high-scale systems is a plus ## Description Lead EarnIn's shift to AI-first reliability engineering. Define how AI transforms on-call, incident response, alert triage, postmortems, and production investigations across SRE and product engineering teams, while setting SLO-driven standards and resilience patterns that enable the company to ship fast and stay safe. The base salary range for this full-time position is $252,000-$308,000, plus equity and benefits. Our salary ranges are determined by role, level, and location. This is a hybrid position in Mountain View (Headquarters) and will require in-office work 2 days a week., * Set a reliability strategy with AI at the center. Define SLIs, SLOs, and error budgets across critical services. Use AI to surface trends, predict capacity risks, and auto-generate reliability scorecards so teams act on data. * Redesign the incident lifecycle around AI-assisted speed. Lead high-severity incident response as IC. Build AI-driven alert correlation and triage that reduces noise and accelerates root-cause identification. Drive adoption of AI-generated postmortems that surface systemic patterns and automatically track corrective actions through to completion. * Improve on-call fundamentally better through automation. Build AI agents that draft runbook responses, pull relevant context from Datadog, incident.io, and Slack during pages, and recommend remediation steps, so on-call engineers spend less time deciding and searching. * Push AI-first operations into product engineering teams. Partner with product engineering to embed AI-assisted investigation, alerting, and production readiness into their workflows. Make AI tooling the default path for every team that owns a service, not an SRE-only capability. * Architect for resilience at scale. Guide service designs for graceful degradation, failure isolation, and capacity planning across EarnIn's AWS footprint (EKS, Kafka, DynamoDB, RDS, SQS). Use AI-driven analysis to identify architectural weak points before they become incidents. * Raise the bar through mentorship and standards. Coach engineers on reliability practices, run design and incident reviews, and build documentation and tooling that makes reliability knowledge accessible. Set the expectation that AI-assisted workflows are how EarnIn operates, not an experiment. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [From Black Box to Glass Box : Bedrock AgentCore Observability](https://www.wearedevelopers.com/videos/2126-from-black-box-to-glass-box-bedrock-agentcore-observability) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [The OpenTelemetry mistakes I keep seeing (and how to stop making them)](https://www.wearedevelopers.com/videos/100158-the-opentelemetry-mistakes-i-keep-seeing-and-how-to-stop-making-them) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [30 powerful AWS hacks in just 30 minutes: Boost your developer productivity](https://www.wearedevelopers.com/videos/1624-30-powerful-aws-hacks-in-just-30-minutes-boost-your-developer-productivity) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)