> Markdown version of [/jobs/ext/131095-senior-site-reliability-engineer-metrics-and-observability](https://www.wearedevelopers.com/jobs/ext/131095-senior-site-reliability-engineer-metrics-and-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - Metrics and Observability - **Company:** CVS Health - **Location:** United States - **Experience:** Expert - **Salary:** $83,430.0 - $203,940.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, BigQuery, Continuous Integration, DevOps, Monitoring of Systems, Reliability Engineering, Power BI, Prometheus, Datadog, Scripting, Google Cloud, Cloud Platform System, Grafana, Reliability of Systems, Uipath, Kubernetes, Appdynamics, Docker, Elk Stack - **Published:** May 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a76790c19e489ccd ## About the Role * 5+ years of experience in Site Reliability Engineering and/or DevOps * 3+ years of experience defining and implementing metrics, SLOs, and SLIs * 2+ years of experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack) * 2+ years of experience with cloud platforms (AWS, Azure, GCP) and container orchestration (Docker, Kubernetes), * Excellent analytical skills and the ability to communicate complex technical concepts to non-technical stakeholders * Familiarity with ITIL and incident management frameworks * Experience with automation tools and scripting languages (e.g., BigQuery, SQL Server, Grafana, DataDog, AppDynamics, OTEL, UiPath, PowerBI) * Knowledge of security best practices in cloud environments, Bachelor's degree or equivalent experience (HS diploma + 4 years relevant experience) ## Description CVS Health Digital is looking for hands-on, passionate people who want to join a high energy and growing team to make a difference in customers' lives and who want to be on the forefront of digital innovation that aims to reinvent what a pharmacy and a health care company can be in the digital world. Currently, we are seeking a dedicated Senior Site Reliability Engineer - Metrics and Observability to focus on metrics, service level objectives (SLOs), service level indicators (SLIs), and error budgets. This individual contributor role will be crucial in enhancing our monitoring and observability practices while taking on automation responsibilities related to quality gates in the release engineering process. The ideal candidate will work closely with cross-functional teams to ensure the reliability and performance of our systems. Expectations for the Role Metrics Development: Define, implement, and maintain key performance metrics, SLOs, and SLIs to measure system reliability and performance. Ensure alignment with business objectives and operational goals. Error Budgets: Manage error budgets effectively, collaborating with development teams to balance reliability and feature delivery. Analyze incidents and outages to inform adjustments to error budgets. Monitoring & Observability: Design and implement comprehensive monitoring solutions to provide real-time visibility into system health. Utilize tools such as Prometheus, Grafana, and other observability platforms to create dashboards and alerts. Quality Gates Automation: Develop and implement automated quality gates that ensure all releases meet defined reliability and performance standards. Collaborate with the release engineering team to integrate these gates into the CI/CD pipeline. Incident Management: Assist in incident response efforts by providing insights from metrics and monitoring tools. Conduct post-mortem analyses to identify root causes and recommend preventive measures. Continuous Improvement: Drive initiatives to enhance monitoring, observability, and reliability practices. Foster a culture of learning and knowledge sharing within the team. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [How to achieve web automation with UiPath](https://www.wearedevelopers.com/videos/310-how-to-achieve-web-automation-with-uipath) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [RPA crash course for .Net developers – intro into the world of RPA from the perspective of a .Net developer](https://www.wearedevelopers.com/videos/271-rpa-crash-course-for-net-developers-intro-into-the-world-of-rpa-from-the-perspective-of-a-net-developer) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)