> Markdown version of [/jobs/ext/1055204-senior-cloud-engineer](https://www.wearedevelopers.com/jobs/ext/1055204-senior-cloud-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Cloud Engineer - **Company:** CVS Health - **Location:** Los Angeles, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $83,430.0 - $222,480.0 - **Contract:** Permanent contract - **Skills:** Cloud Computing, Cloud Computing Security, Cloud Engineering, Apache Lucene, Query Languages, Elasticsearch, Fault Tolerance, Open Source Technology, Migration Manager, Datadog, Cloud Platform System, Grafana, Technical Debt, Infrastructure as Code (IaC), BIG-IP Access Policy Manager (APM), New Relic (SaaS), Data Pipelines, Dynatrace - **Published:** June 30, 2026 - **Apply:** https://dejobs.org/x/x/EDF7DEA877874EEAB5AA5A96601720DA/job/ ## About the Role * 5 years of Demonstrated expertise on managing all aspects of a Grafana Cloud installation. * 2 years of experience managing Infrastructure as Code (IaC) or related development experience. * Experience as a technical leader. Preferred Qualifications: * Proven Migration Experience: Technical leadership in at least one large-scale migration away from proprietary APM platforms (specifically New Relic or Datadog) or heavy log-analytics platforms (Elasticsearch/ELK) into a unified open-source or Grafana-centric stack. * Deep Query Language Fluency: Expert-level mastery of PromQL, LogQL, and TraceQL, with a strong foundational understanding of how to map them from legacy languages like NRQL or Lucene syntax. * Strategic Influence: Ability to build a coalition of support across the organization. You are a respected thinker who can frame technical decisions (rationales, benefits, risks) in business terms to stakeholders. * Communication: Effective at sharing information across audiences (from direct leaders to engineering communities). Skilled at tailoring information and key points to different levels of technical and non-technical stakeholders. * Execution: Adept at planning, delivering, and supporting complex projects. Maintains composure in stressful or unexpected situations and can pivot action plans without affecting the broader process. * Growth Mindset: Committed to developing yourself and others. Proactively focuses on professional development to accomplish new tasks and team goals., * Bachelor's degree preferred/specialized training/relevant professional qualification. ## Description We are currently unified on Grafana as our target platform and are in the midst of a strategic, high-impact migration away from New Relic and Elastic. As the Observability Lead, you will spearhead this modernization-decommissioning legacy telemetry footprints, unifying disparate data models, and establishing Grafana as our singular, enterprise-wide pane of glass. You will participate in the design, implementation, and maintenance of cloud-based infrastructure, ensuring our architecture is scalable, secure, and resilient.Key Responsibilities, * Lead the Modernization Wave: Architect and execute a seamless migration strategy to transition legacy APM, distributed tracing, and log management workloads off New Relic and Elastic onto a modern Grafana (LGTM) and OpenTelemetry architecture. * Consolidate & Standardize Data Models: Translate legacy NRQL/Lucene querying paradigms and proprietary formats into scalable PromQL, LogQL, and TraceQL structures, ensuring no loss of historical operational context during the cutover. * Optimize Telemetry Architecture: Audit existing data pipelines to reduce redundancy, manage label cardinality, and implement aggressive data-filtering strategies that maximize Grafana Mimir/Loki efficiency while driving down total cost of ownership. * Establish Operational Excellence: Build and refine the right metrics for observability; foster a culture of operational excellence where telemetry data drives product and service success. Cloud Leadership & Engineering * Strategic Architecture: Define future-state cloud platform strategies. Architect and design for resiliency, fault tolerance, and self-healing. * Governance & Standardization: Enforce cloud security policies, schema protocols, and best practices. Establish engineering standards that enhance efficiency and maintainability across multiple products. * Influencing the Roadmap: Influence engineering roadmaps (e.g., tech debt, NFRs) and proactively identify technical gaps, risks, and mitigations. Participate in leadership meetings to align technical outcomes with business value. * Coaching & Mentorship: Provide mentorship and guidance to junior through senior-level engineers. Take a hands-on approach to skill development, fostering growth in cloud engineering domains. * Cross-Functional Collaboration: Act as a go-to resource for complex technical challenges. Collaborate inclusively across teams, harmonizing differing opinions and steering discussions toward productive outcomes. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)