> Markdown version of [/jobs/ext/2689801-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2689801-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability engineer - **Company:** Filevine, Inc. - **Location:** United States - **Experience:** Expert - **Salary:** $175,000.0 - $195,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Bash Shell, Continuous Integration, DevOps, Disaster Recovery, Python (Programming Language), Reliability Engineering, Newrelic, Datadog, Data Logging, Grafana, Kubernetes, Machine Learning Operations - **Published:** September 3, 2026 - **Apply:** https://www.dice.com/job-detail/5d481ba0-58d2-4e3e-87f7-dad6c84d7678 ## About the Role * Experience: 8+ years in software/platform engineering or DevOps, including 5+ years dedicated to Site Reliability Engineering. * Infrastructure & Observability: Expertise in cloud platforms (AWS), Kubernetes, IaC, and full-stack observability (tracing, logging, SLI/SLOs). Grafana, Datadog, NewRelic * Automation & Scripting: Proficient in Python, Go, or Bash for building CI/CD pipelines, production tools, and toil-reducing automation. * Incident Leadership: Proven track record in root cause analysis, high-severity incident response, and long-term reliability engineering. * Leadership & Mentorship: Strong communication skills with a history of mentoring engineers, leading cross-functional projects, and setting technical strategy. * Operational AI/ML: Hands-on experience using AI/ML on telemetry data to predict capacity risks, spot anomalies, and optimize system workflows. ## Description * Observability & Alerting: Design and improve monitoring, logging, tracing, dashboards, and SLI/SLOs for production visibility. * Automation & CI/CD: Build internal tools and delivery pipelines to boost efficiency, eliminate toil, and ensure reliable deployments. * Reliability & System Quality: Drive continuous improvements in system performance, scalability, and security to mitigate customer impact. * Incident Management: Lead production incident response from triage to resolution, turning lessons into durable runbooks and preventatives. * Technical Leadership: Guide major technical initiatives, align cross-team engineering efforts, and manage technical risks. * Mentorship: Elevate SRE team capability through design reviews, paired problem-solving, and incident post-mortems. * Operations & On-Call: Join the on-call rotation while leading capacity planning and disaster recovery readiness. * AI/ML Operationalization: Leverage operational AI/ML tools to forecast capacity risks, detect patterns, and automate system health. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)