> Markdown version of [/jobs/ext/1469175-manager-devops-automation-google-cloud-platform](https://www.wearedevelopers.com/jobs/ext/1469175-manager-devops-automation-google-cloud-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Manager (DevOps, Automation, Google Cloud Platform) - **Company:** CVS Health - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $106,605.0 - $260,590.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Application Release Automation, Microsoft Azure, Software as a Service, Cloud Computing, Cloud Engineering, Cluster Analysis, DevOps, Distributed Systems, Fault Tolerance, Github, Monitoring of Systems, Information Technology Operations, Python (Programming Language), Machine Learning, Reliability Engineering, Site Reliability Engineering Practices, Prometheus, Software Engineering, Software Vulnerability Management, Datadog, Diagnostic Tools, Google Cloud, Cloud Platform System, System Availability, Delivery Pipeline, Grafana, Mttr, Reliability of Systems, Infrastructure as Code (IaC), Cloudformation, Containerization, Kubernetes, Information Technology, Performance Monitor, ArcSight Event Correlation, Terraform, Splunk, Appdynamics, Dynatrace, Devsecops, Docker, Jenkins, Vulnerability Analysis, Golang, Microservices - **Published:** July 28, 2026 - **Apply:** https://www.dice.com/job-detail/bdca35a5-ad38-49e4-bdae-b1aadc507eef ## About the Role * 5+ years of experience in software engineering, SRE, or production engineering within large-scale distributed systems. * Hands-on experience with AIOps or intelligent monitoring platforms, including anomaly detection and event correlation. * Experience with observability tools such as AppDynamics, Grafana, Prometheus, and Splunk. * Strong expertise in cloud platforms (AWS, Azure, or Google Cloud Platform), cloud-native architectures (Kubernetes, containers, microservices), and CI/CD pipelines (GitHub Actions, Jenkins). * Experience with Infrastructure as Code (Terraform, AWS CloudFormation, or Google Cloud Platform Deployment Manager). * Proficiency in at least one programming language (e.g., Python, Java, Go). * Proven track record of improving operational metrics (SLAs, CSAT, resolution time). * Experience with GenAI and automation tools such as OpenAI, Copilot, Gemini, Claude, and MCP. * Strong understanding of distributed systems, resiliency patterns, and fault tolerance. * Strong analytical skills with experience in reporting and performance measurement. * Excellent communication, stakeholder management, and conflict resolution skills. * Ability to thrive in a fast-paced, high-growth, or matrixed environment., * Experience designing and implementing AIOps platforms or predictive reliability systems at scale. * Strong knowledge of machine learning applications in IT operations (e.g., anomaly detection, forecasting, clustering). * Experience defining and managing SLIs/SLOs and error budgets at scale. * Experience with OpenTelemetry and modern observability standards. * Familiarity with chaos engineering, resilience testing, and fault injection frameworks. * Exposure to GenAI-driven operations or AI-assisted troubleshooting tools. * Experience in healthcare, financial services, enterprise SaaS, or other regulated industries. * Proven ability to lead cross-functional initiatives and influence senior stakeholders. * Contributions to open-source projects related to SRE, observability, or AIOps. * Relevant certifications in cloud platforms, SRE, AIOps, OpenTelemetry, or DevOps are a plus. Education: * Bachelor's degree (or equivalent experience) in Computer Science, Engineering, or a related discipline. Leadership Competencies: * Customer-first mindset with a passion for delivering exceptional service. * Strategic thinker with strong execution capabilities. * High emotional intelligence with strong people leadership skills. * Continuous improvement mindset with a focus on innovation. ## Description Join Fortune 7 CVS Health as a Sr. Manager, Software Engineering - DevOps, Observability & Monitoring to lead strategic initiatives for the CVS Caremark Digital team. In this role, you will drive the vision and execution of modern platform engineering capabilities that enable scalable, reliable, and secure application development. You will build and lead a high-performing engineering team responsible for designing, implementing, and operating cloud-native platforms. The ideal candidate is a hands-on leader with deep technical expertise in cloud architectures, DevOps, and observability, combined with the ability to lead transformation and operational excellence initiatives., Leadership & People Management: * Lead, mentor, and grow a team of software engineers and SRE/DevOps engineers. * Foster a culture of accountability, innovation, and continuous improvement. * Define team goals, OKRs, and performance metrics aligned with organizational strategy. * Partner with product, architecture, and business stakeholders to deliver platform capabilities. AIOps & Intelligent Operations: * Drive the adoption of AIOps solutions for predictive monitoring, anomaly detection, and automated root cause analysis. * Integrate machine learning models and analytics into monitoring pipelines to proactively detect and prevent incidents. * Develop intelligent alerting systems to reduce noise and improve signal quality. Observability & Monitoring: * Architect and implement scalable observability frameworks across metrics, logs, traces, and events. * Establish standards for instrumentation, telemetry collection, and distributed tracing. * Enable proactive monitoring, alerting, and incident detection using tools such as Datadog, Prometheus, Grafana, and Splunk. * Define and implement enterprise-wide SRE practices, including SLIs, SLOs, error budgets, and reliability governance. DevOps & Platform Engineering: * Drive adoption of CI/CD pipelines, Infrastructure as Code (IaC), and GitOps practices. * Lead the design and evolution of scalable, automated, and secure platform engineering solutions. * Standardize development and deployment workflows across teams. * Champion DevOps maturity, developer productivity, and release automation. Reliability & Incident Management: * Improve system reliability through error budgets, resiliency patterns, and chaos engineering. * Lead incident response processes, postmortems, and root cause analysis. * Drive continuous improvements in MTTR (Mean Time to Recovery) and system availability. Cloud & Infrastructure: * Architect and manage solutions across cloud platforms (AWS, Azure, or Google Cloud Platform). * Ensure scalability, security, and cost optimization of infrastructure. * Oversee containerization and orchestration using Docker and Kubernetes. Security & Compliance: * Integrate DevSecOps practices into pipelines and platform tooling. * Ensure compliance with enterprise security standards and regulatory requirements. * Automate security scanning, vulnerability management, and policy enforcement. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [What’s the Difference between a Junior, Mid, and Senior Developer?](https://www.wearedevelopers.com/magazine/238-what-s-the-difference-between-a-junior-mid-and-senior-developer)