Data Engineer

ICF Consulting Group, Inc.
Atlanta, GA, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$98,614.0 - $167,644.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Big Data Bioinformatics Cloud Database Cloud Engineering Cloud Storage Databases Information Engineering Data Integration Data Systems Apache Hive
+21 more
Python (Programming Language) NumPy SAS (Software) SQL Databases Subversion XPath Extensible Stylesheet Language Transformations (XSLT) Data Processing Google Cloud Cloud Platform System Git Pandas Matplotlib Data Lakes Pyspark Information Technology Plotly Tools for Reporting Software Version Control Data Pipelines Serverless Computing

Job description

ICF is seeking a detail-oriented Data Engineer to support the development, optimization, and deployment of large-scale healthcare data pipelines and analytical datasets. This role will focus on cancer registry and population-based data systems. As part of the team contributing to the fight against cancer, you will work with high-volume, complex structured datasets using NAACCR registry-standard formats. ICF partners with the Centers for Disease Control and Prevention (CDC) to monitor cancer diagnoses, incidence rates, geographic distribution, and treatment patterns.

These efforts enable researchers, clinicians, and policymakers to better understand and combat cancer. In this role, you may also have the opportunity to support a diverse range of projects across multiple clients and subject areas, allowing you to expand your technical expertise and gain new skills. Historically, most of our projects are small to medium in size, providing team members with the opportunity to play an active role throughout the full project lifecycle-from requirements gathering and design to implementation, validation, and maintenance.

This position can be based in Rockville, MD or Atlanta, GA and offers a hybrid work arrangement, with 2-3 days per week onsite at the ICF office., * Design, develop, and maintain scalable, cloud-based data pipelines using Python, PySpark, and SQL in cloud environments

  • Process and standardize large-scale healthcare datasets, including cancer registry, mortality, and population data
  • Develop and maintain production databases, analytic datasets, and reporting tables for statistical analysis and surveillance
  • Perform rigorous data quality validation, reconciliation, and consistency checks across large datasets
  • Collaborate with epidemiologists, statisticians, and analysts to design and implement data solutions for analytics needs
  • Troubleshoot data pipeline failures and large-scale data processing issues

Requirements

  • Position requires a minimum of 2 years of experience in data engineering or working with healthcare data environments
  • BA/BS in computer science, Data Science, Statistics, Public Health Informatics, Bioinformatics, or related field
  • The position requires strong expertise in data integration, transformation, quality control, and the delivery of production-ready data for analytics and reporting platforms.
  • Demonstrated hands-on experience with Python, PySpark, Spark SQL, and commonly used Python libraries for data processing and analysis
  • Experience with Python libraries such as pandas, NumPy, PyArrow, matplotlib, Plotly, or similar tools
  • Position requires experience with cloud platforms such as AWS, Azure, Google Cloud, or similar cloud environments
  • Experience handling large databases and high-volume datasets (millions+ records)
  • Must have experience with SAS and/or R
  • Experience implementing data quality validation and QC frameworks
  • Strong analytical, troubleshooting, and problem-solving skills
  • Must have the ability to obtain and maintain a Public Trust

Preferred Skills/Experience:

  • Experience with XML/XSLT/XML path
  • Proven experience migrating SAS-based pipelines to Python/PySpark or cloud-native architectures
  • Experience with cloud-based data platforms, data lakes, cloud storage, serverless functions, and containerized applications.
  • Experience developing and optimizing PySpark pipelines for large-scale data processing
  • Version control with (Example: Git, SVN) Professional Experience:
  • Works well in a team environment and values the expertise and of others
  • Understands the value of processes and protocol, and is willing to follow them
  • Excellent written and verbal communication skills
  • Strong analytical and problem-solving skills

About the company

ICF is a global advisory and technology services provider, but we’re not your typical consultants. We combine unmatched expertise with cutting-edge technology to help clients solve their most complex challenges, navigate change, and shape the future.

We can only solve the world’s toughest challenges by building a workplace that allows everyone to thrive. We are an equal opportunity employer. Together, our employees are empowered to share their expertise and collaborate with others to achieve personal and professional goals. For more information, please read our EEO policy.

We will consider for employment qualified applicants with arrest and conviction records.

Reasonable Accommodations are available, including, but not limited to, for disabled veterans, individuals with disabilities, and individuals with sincerely held religious beliefs, in all phases of the application and employment process. To request an accommodation, please email and we will be happy to assist. All information you provide will be kept confidential and will be used only to the extent required to provide needed reasonable accommodations.

Read more about workplace discrimination rights or our benefit offerings which are included in the Transparency in (Benefits) Coverage Act.

Candidate AI Usage Policy

At ICF, we are committed to ensuring a fair interview process for all candidates based on their own skills and knowledge. As part of this commitment, the use of artificial intelligence (AI) tools to generate or assist with responses during interviews (whether in-person or virtual) is not permitted. This policy is in place to maintain the integrity and authenticity of the interview process.

However, we understand that some candidates may require accommodation that involves the use of AI. If such an accommodation is needed, candidates are instructed to contact us in advance at We are dedicated to providing the necessary support to ensure that all candidates have an equal opportunity to succeed.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou Ā· Coffee With Developers

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell Ā· LIVE

3:18 min

Identifying deceptive server content and parsing messy HTML

Vidas Bacevičius Vidas Bacevičius Ā· WWC 2025

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer Ā· Coffee With Developers

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon Ā· WWC Europe 2026

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham Ā· WWC 2025

Videos

See all

Related articles

See all