> Markdown version of [/jobs/ext/2705458-data-engineer](https://www.wearedevelopers.com/jobs/ext/2705458-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** CVS Health - **Location:** New York, TX, United States - **Experience:** Expert - **Salary:** $92,700.0 - $222,480.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Agile Methodology, Artificial Intelligence, Airflow, Business Analytics Applications, Bash Shell, Big Data, BigQuery, Unix, Cloud Computing, Cloud Engineering, Cloud Storage, Software Quality, Information Systems, Continuous Integration, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Security, Dataspaces, Data Structures, Data Warehousing, DevOps, Data Flow Control, Github, Identity and Access Management, Java Web Services, Python (Programming Language), Machine Learning, Metadata, NoSQL, Query Optimization, Cloud Services, Cloudera, Service-Oriented Architecture, SQL Databases, Data Streaming, Systems Integration, Unix Commands, Unstructured Data, Data Logging, Data Processing, Scripting, Google Cloud, Cloud Platform System, Cloud Monitoring, Large Language Models, Generative AI, Build Server, Git, Data Lakes, Kubernetes, Information Technology, Google Cloud Functions, Data Analytics, Apache Kafka, Spark Streaming, Data Management, Api Design, GPT, Software Version Control, Data Pipelines, Jenkins, Programming Languages, Microservices - **Published:** September 4, 2026 - **Apply:** https://www.dice.com/job-detail/c1bec1fb-6295-43d9-a1da-f7a1e2d8c147 ## About the Role * 3-5+ years of experience with SQL, NoSQL * 3-5+ years of experience with Python (or a comparable scripting language) * 3+ years of experience with Data warehouses (such as data modeling and technical architectures) and infrastructure components * 3+ years of experience with ETL/ELT, and building high-volume data pipelines * 3+ years of experience with reporting/analytic tools * 3+ years of experience with Query optimization, data structures, transformation, metadata, dependency, and workload management * 3+ years of experience with Big data and cloud architecture * 3+ years of hands-on experience building and managing cloud infrastructure on Google Cloud Platform (Google Cloud Platform), including Cloud Storage, BigQuery, Dataflow, and Dataproc * 3+ years of experience deploying and scaling containerized applications on Google Cloud Platform using Google Kubernetes Engine (GKE), Cloud Run, and Artifact Registry * 3+ years of experience with real-time and streaming data technologies on Google Cloud Platform (i.e. Pub/Sub, Dataflow, Cloud Functions, Apache Kafka, Spark Streaming) * 1+ year(s) of soliciting complex requirements and managing relationships with key stakeholders * 1+ year(s) of experience independently managing deliverables Preferred Qualifications * Experience in designing and building data engineering solutions in cloud environments (preferably Google Cloud Platform) * Experience with Git, CI/CD pipeline, and other DevOps principles/best practices * Experience with bash shell scripts, UNIX utilities & UNIX Commands * Ability to leverage multiple tools and programming languages to analyze and manipulate data sets from disparate data sources * Knowledge of API development * Exposure to Generative AI * Exposure Large Language Models (GPT, Claude, Gemini, Llama, etc.) * Experience with complex systems and solving challenging analytical problems * Strong collaboration and communication skills within and across teams * Knowledge of data visualization and reporting * Experience with schema design and dimensional data modeling * Google Professional Data Engineer Certification * Knowledge of microservices and SOA * Formal SAFe and/or agile experience. Previous healthcare experience and domain knowledge * Experience designing, building, and maintaining data processing systems * Experience architecting and building data warehouse and data lakes Education Bachelor's Degree or equivalent work experience in Computer Science, Information Systems, Data Engineering, Data Analytics, Machine Learning, or related field required. Master's Degree preferred. ## Description This is a hybrid position. Looking for candidates local to MA, CT, NY, TX If you're eager to make a real impact in the health care industry through your own meaningful contributions, explore a role in technology with CVS Health. Our journey calls for technical innovators and data visionaries: come help us pave the way. At CVS Health, we possess an extensive repository of healthcare data that spans over 150 million individuals, providing an unparalleled foundation for ambitious Data Engineers. In this role, you will engage with complex business challenges, harnessing modern tools and technologies to securely store, process, transform, and enrich terabyte to petabyte scale healthcare data. Your work will underpin data-driven business decisions and contribute to our mission of delivering industry-best data products / software with a customer-first mindset and team-oriented approach. As a Senior Data Engineer, you will be instrumental in designing, developing, and maintaining optimal data pipelines to assemble large and intricate datasets, catering to the business requirements of various CVS lines of business. Collaborating closely with teams, you will craft tools to provide actionable insights and integrate them with consumer touchpoints. In this role, you will: * Architect and develop robust, scalable ETL/ELT pipelines using Cloud Dataflow, Cloud composer (Airflow), and Pub/Sub for both batch and streaming use cases. Leverage BigQuery as the central data warehouse and design integrations with other Google Cloud Platform services (e.g., Cloud storage, Cloud functions). * Build and optimize analytical data models in BigQuery. Implement partitioning, clustering, and materialized views for performance and cost efficiency. Ensure compliance with data governance, access controls, and IAM best practices. * Develop integrations with external systems (APIs, flat files etc.) using Google Cloud Platform-native or hybrid approaches. Utilize tools like Dataflow or custom Python/Java services on Cloud Functions or Cloud Run to handle transformations and ingestion logic. * Build automated CI/CD pipeline using Cloud Build, GitHub Actions, or Jenkins for deploying data pipeline code and workflows. Set up observability using Cloud Monitoring, Cloud Logging, and Error Reporting to ensure pipeline reliability. * Lead architectural decisions for data platforms and mentor junior engineers on cloud-native data engineering patterns. Promote best practices for code quality, version control, cost optimization, and data security in a Google Cloud Platform environment. Drive initiatives around data democratization, including building reusable datasets and data catalogs via Dataplex or Data Catalog. * Design, develop, and maintain enterprise AI/ML solutions and platforms that address complex business and customer challenges. Integrate Large Language Models (LLMs), generative AI capabilities, and AI agents into products and data ecosystems. Define the technical architecture and infrastructure required for AI applications, and build scalable platforms supporting model training, deployment, monitoring, governance, and lifecycle management while ensuring security, reliability, and operational excellence. As leaders in healthcare, our analytics and engineering teams deliver innovative solutions to business problems by collaborating with cross-functional teams in a dynamic and agile environment. You will be part of a team that values collaboration and encourages innovative thinking at all levels. You will be intellectually challenged to solve problems associated with large scale complex, structured and unstructured data, that will allow you to grow your technical skills and engineering expertise. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [WeAreDevelopers LIVE - Node and Package Security](https://www.wearedevelopers.com/videos/2138-wearedevelopers-live-node-and-package-security) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Streaming AI Responses in Real-Time with SSE in Next.js & NestJS](https://www.wearedevelopers.com/videos/1630-streaming-ai-responses-in-real-time-with-sse-in-next-js-nestjs) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)