> Markdown version of [/jobs/ext/2462737-lead-data-engineer-remote-wfh](https://www.wearedevelopers.com/jobs/ext/2462737-lead-data-engineer-remote-wfh). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Engineer - Remote (WFH) - **Company:** Cognitive Medical Systems - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $130,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Airflow, Amazon Web Services, Amazon S3, Unit Testing, Bash Shell, Big Data, CentOS, Data Transmissions, Information Engineering, Extract Transform Load (ETL), Data Warehousing, Database Queries, Github, Python (Programming Language), Linux Kernel, Linux System Administration, Red Hat Enterprise Linux, Reference Data, Shell Script, SonarQube, Data Processing, Data Ingestion, Snowflake, Sonatype, Pyspark, Information Technology, Functional Programming, Amazon Simple Queue Service (SQS), Software Version Control, Data Pipelines, Serverless Computing, Jenkins, Databricks, Artifactory - **Published:** August 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=559af16c01f48ef7 ## About the Role * Bachelor's degree in Computer Science, Data Engineering, or related field; 8 or more years in data engineering, including 4 or more years building and operating high-volume batch ETL systems. * Expert-level Python andPySpark * Production experience with Apache Airflow (AWS MWAA) for pipeline orchestration at scale. * Production experience with AWS Redshift Serverless, Snowflake, and Amazon Athena for data warehousing and query at scale. * Hands-on experience with Databricks (Notebooks, Jobs) for large-scale data processing. * Strong SQL skills for multi-source data transformation, editing logic, and reconciliation of data pipelines. * Experience with AWS S3, SQS, Lambda, and Event Bridge for data ingestion and event-driven orchestration. * Experience with Linux environments including RHEL, CentOS, and Amazon Linux 2, and shell scripting in Bash. * Demonstrated experience operating data pipelines under strict data quality, accuracy, and timeliness SLAs in a federal or healthcare environment. * Experience with version control and CI/CD integration using GitHub, Jenkins, and JFrog Artifactory. * Ability to pass CMS and internal required background checks for public trust. ## Description *This position is contingent upon contract award* The Lead Data Engineer is responsible for PDE processing pipelines, data quality, and IDR integration for the Drug Data Processing System (DDPS) and Payment Reconciliation System (PRS) O&M contract. This individual owns the end-to-end PDE data pipeline - receiving, validating, editing, and storing approximately 9 to 10 million Prescription Drug Event records daily - and maintains all vendor reference data integrations including FDA, NCPDP, FDB, MediSpan, NPPES, and OIG data sources. This is a remote position; however, Cognitive hires only in the following designated U.S. states based on contract and business requirements: VA, DC, MD, TN, FL, AZ, CO, OR, and TX., * Own summary records of prescription drug transactions, also known as pharmacy drug events (PDE), ingestion, validation, editing, and storage pipelines. * Apply and maintain CMS-defined business rule validations, returning edit results to plan sponsors, and minimizing PDEs requiring manual analysis. * Engineer and maintain all vendor reference data edits from a variety of sources for drug event record, pharmacy data and files, exclusion and preclusion list; provide health plan and CMS status updates on PDE submissions. * Build, maintain, and optimize Apache Airflow and PySpark transformations across variety of technologies, including S3, Redshift Serverless, Snowflake, and Databricks. * Maintain system compatibility with other CMS data model changes. * Monitor PDE submissions for adjustment trends and informational edit rates; provide analytical capability to link rejected PDEs with IDR records; provide beneficiary and plan-level cost aggregations to PRS for year-end reconciliation. * Support data transfer projects and new data share integrations with CMS and downstream stakeholders including API and multi-cloud approaches as directed by the CMS. * Maintain target unit test coverage for all new pipeline code; support BDD and TDD practices; track and surface real-time PDE processing metrics including volumes, error rates, and reconciliation accuracy. * Operate within the CMS Lean-Agile Release Train adhering to the Iteration Schedule; deliver code through GitHub, Jenkins, JFrog Artifactory/XRay, SonarQube, and Snyk CI/CD pipeline. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Best Paying Remote Jobs](https://www.wearedevelopers.com/magazine/255-best-paying-remote-jobs)