> Markdown version of [/jobs/ext/1216986-staff-observability-platform-engineer-sre](https://www.wearedevelopers.com/jobs/ext/1216986-staff-observability-platform-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Observability Platform Engineer (SRE) - **Company:** CVS Health - **Location:** Richardson, TX, United States - **Experience:** Experienced - **Salary:** $118,450.0 - $236,900.0 - **Contract:** Temporary contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Microsoft Azure, BigTable, Business Software, Cloud Computing, Cloud Engineering, Databases, Continuous Integration, Relational Databases, DevOps, Monitoring of Systems, Python (Programming Language), PostgreSQL, MySQL, NoSQL, Octopus Deploy, Online Transaction Processing, Release Management, Reliability Engineering, Prometheus, Software Engineering, Data Streaming, Systems Integration, Data Logging, Cloud Platform System, Computer Network Technologies, Istio, System Availability, Large Language Models, Grafana, Reliability of Systems, Backend, Cloudformation, Kubernetes, Infrastructure Automation Frameworks, Low Latency, Cassandra, Apache Kafka, Data Management, Vertica, Terraform, Splunk, Appdynamics, Data Pipelines, Docker - **Published:** July 9, 2026 - **Apply:** https://www.juju.com/job/00000000gessoh ## About the Role + 10+ years of experience in Software Engineering, Platform Engineering, or SRE. + 7+ years of experience with observability practices, including SLIs/SLOs/SLAs, alerting, and incident management. + 7+ years building production-grade backend services in Java/python. + 7+ years implementing and operating OpenTelemetry, including OTLP, semantic conventions, and instrumentation patterns. + 7+ years with cloud-native and containerized platforms (Docker, Kubernetes, Argo CD). + 7+ years working with public cloud platforms (AWS, GCP, or Azure). + 5+ years designing and scaling distributed, high-volume data pipelines. + 5+ years working with Grafana OSS or comparable observability backends (e.g., Grafana, Loki, Tempo, Prometheus). + 5+ years with relational databases (PostgreSQL, MySQL). PREFERRED QUALIFICATIONS + Excellent analytical skills and the ability to communicate complex technical concepts to non-technical stakeholders + Experience with service meshes and networking technologies such as Envoy and Istio + Experience integrating or operating commercial observability platforms (Splunk, AppDynamics, etc.) + Experience with streaming and data platforms such as Kafka, Pulsar, or similar technologies + Familiarity with time-series, NoSQL, or analytical databases (ClickHouse, Bigtable, Cassandra, etc.) + Experience with Infrastructure as Code tools such as Terraform or CloudFormation + Experience with cost optimization and capacity planning for large-scale cloud infra + Experience with chaos engineering, resiliency testing, or fault injection + Background in security-aware platform design, including secure service-to-service communication + Experience mentoring senior engineers and influencing platform standards across organizations + Strong operational experience supporting 24x7 production systems, including on-call responsibilities + Knowledge of security best practices in cloud environments, Bachelor's degree or equivalent experience (HS diploma + 4 years relevant experience) ## Description CVS Health PBM is looking for hands-on, passionate people who want to join a high energy and growing team, who want to be on the forefront of digital innovation that aims to reinvent what a pharmacy and a health care company can be in the digital world. As a **Lead Platform Reliability Engineer** , you will design and implement metrics and observability frameworks with a strong focus on service level objectives (SLOs), service level indicators (SLIs), error budgets, and cloud infrastructure scaling and capacity estimation. This individual contributor role is critical to enhancing our monitoring and observability capabilities, while also driving automation initiatives related to quality gates within the release engineering process. You will work closely with cross-functional teams to ensure the reliability, performance, and scalable growth of our cloud-based systems. _Expectations for the Role:_ **Metrics Development:** Define, implement, and maintain key performance metrics, SLOs, and SLIs to measure system reliability and performance. Ensure alignment with business objectives and operational goals. **Error Budgets:** Manage error budgets effectively, collaborating with development teams to balance reliability and feature delivery. Analyze incidents and outages to inform adjustments to error budgets. **Monitoring & Observability:** Design and implement comprehensive monitoring solutions to provide real-time visibility into system health. Utilize tools such as Prometheus, Grafana, Loki, Temp and other observability platforms to create dashboards and alerts. **Cloud Infrastructure Scaling:** Architect, design, and implement scalable cloud infrastructure capable of supporting multiple business applications, ensuring reliability, performance, and future growth. **Quality Gates Automation:** Develop and implement automated quality gates that ensure all releases meet defined reliability and performance standards. Lead the release Devops team to integrate these gates into the CI/CD pipeline. **Incident Management:** Assist in incident response efforts by providing insights from metrics and monitoring tools. Conduct post-mortem analyses to identify root causes and recommend preventive measures. **AIOps Insight Automation:** Use AI to surface **what changed / what's abnormal / next best action** from metrics, logs, and traces-minimizing manual dashboard analysis. **AI-Accelerated Incident Response:** Apply GenAI to speed **triage and RCA** with fast signal summarization and guided investigation paths. **AI/LLM Observability & Governance:** Monitor AI workloads for **quality, safety, cost, latency, reliability** with **end-to-end tracing (request * prompt * tools * output)** and secure logging/redaction. **AI-Backed Release Quality Gates:** Embed AI signal checks into CI/CD to flag **SLO risk, latency/error drift, and regression patterns** before production release. ## Related Videos - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)