> Markdown version of [/jobs/ext/3298844-site-reliability-engineer-retail-pharmacy](https://www.wearedevelopers.com/jobs/ext/3298844-site-reliability-engineer-retail-pharmacy). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - Retail Pharmacy - **Company:** CVS Health - **Location:** Scottsdale, AZ, United States - **Experience:** Experienced - **Salary:** $92,700.0 - $203,940.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), JavaScript (Programming Language), Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Software Code Optimization, DevOps, Distributed Systems, Fault Tolerance, Github, Monitoring of Systems, Python (Programming Language), OpenShift, Performance Tuning, Windows PowerShell, Systems Development Life Cycle, Reliability Engineering, Prometheus, Software Engineering, Datadog, Google Cloud, Cloud Platform System, Delivery Pipeline, Grafana, Mttr, HybridCloud, Gitlab, Containerization, Kubernetes, Information Technology, Deployment Automation, Rancher, Performance Monitor, Bitbucket, Splunk, Dynatrace, Docker, Jenkins, Golang, Microservices - **Published:** September 16, 2026 - **Apply:** https://dejobs.org/x/x/5A31870356444BFF986FF01884AF2FED/job/ ## About the Role * 5+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Infrastructure Engineering, or related technology roles. * 3+ years of experience delivering and supporting large-scale distributed systems utilizing reliability and resiliency concepts. * 2+ years of experience with one or more programming languages such as Java, Python, Go, or JavaScript. * 2+ years of experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform. * Hands-on experience with Kubernetes, OpenShift, Docker, Rancher, and containerized workloads. * Experience implementing and supporting CI/CD pipelines using tools such as GitHub, Bitbucket, Jenkins, GitLab, or similar platforms. * Experience with observability and monitoring tools such as Splunk, Dynatrace, Datadog, Prometheus, Grafana, OpenTelemetry, or similar technologies. * Strong scripting and automation skills using Shell, Python, PowerShell, or equivalent technologies. * Experience supporting microservices-based and cloud-native architectures. * Working knowledge of Incident Management, Problem Management, Change Management, and ITIL-based operational practices. * Excellent analytical, troubleshooting, communication, and collaboration skills., * Experience supporting retail, pharmacy, healthcare, or large-scale edge computing environments. * Experience designing and implementing SLO, SLA, and Error Budget frameworks. * Knowledge of distributed tracing and observability best practices. * Experience driving platform modernization, reliability engineering initiatives, and operational excellence programs. * Certifications such as AWS Certified Solutions Architect, Kubernetes (CKA/CKAD), Google SRE, or related cloud certifications., Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience. ## Description The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, scalability, and performance of CVS Health's Retail and Pharmacy platforms. This role combines software engineering, operations, observability, and automation practices to proactively identify and resolve issues, improve system resilience, and support critical store operations. As part of the SRE organization, you will partner with application development, infrastructure, observability, and store operations teams to drive operational excellence, implement reliability engineering best practices, and enable highly scalable deployments across thousands of retail and pharmacy locations., Observability & Monitoring * Develop and implement proactive monitoring, alerting, and dashboarding strategies to detect issues before they impact store operations or customer experience. * Design and maintain operational dashboards using enterprise observability platforms. * Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budgets for critical business services. * Analyze platform telemetry, logs, traces, and metrics to improve service reliability and reduce operational risk. * Drive continuous improvements in observability maturity across Edge applications and services. Reliability Engineering & Incident Management * Lead major incident response, recovery, and post-incident reviews to minimize customer impact and prevent recurring issues. * Perform root cause analysis and drive corrective and preventive actions through structured Problem Management practices. * Improve key operational metrics including Mean Time to Detect (MTTD), Mean Time to Resolve (MTTR), and service availability. * Collaborate with engineering teams to build reliability into applications throughout the Software Development Lifecycle (SDLC). * Drive automation initiatives to reduce operational toil and improve system resiliency. Performance & Platform Optimization * Identify and eliminate bottlenecks in development, testing, and deployment workflows. * Support performance tuning and capacity planning for Edge retail and pharmacy applications. * Analyze system behavior and implement improvements that enhance scalability, stability, and efficiency. * Partner with infrastructure teams to maintain highly available and resilient platform services. Edge Platform Operations * Support business-critical applications deployed across CVS retail and pharmacy locations. * Collaborate with store operations and engineering teams to ensure seamless operation of Edge platforms. * Participate in on-call rotations and provide technical leadership during production incidents. * Ensure operational readiness, deployment validation, and production support for new platform capabilities. Cloud, Microservices & Deployment Engineering * Champion cloud-native technologies and container-based architectures. * Support and optimize Kubernetes and OpenShift environments operating in hybrid cloud ecosystems. * Leverage CI/CD pipelines and Infrastructure-as-Code principles to enable automated, scalable deployments. * Promote best practices for microservices architecture, resiliency, and deployment automation. * Work closely with development teams to ensure services are observable, scalable, and production ready. ## Related Videos - [Reducing Cognitive Overload Through Platform Engineering](https://www.wearedevelopers.com/videos/679-reducing-cognitive-overload-through-platform-engineering) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)