Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
We are hiring a Data Engineer to build a lightweight, non-production data environment that mirrors the functionality and core data structures of a partner environment at a smaller scale. This person will stand up the AWS-based environment in partnership with IT, configure Databricks and QuickSight, create an initial data lake, and progressively develop warehouse schemas and tables for analytics and AI experimentation. The role will ingest data from flat files, Excel, and other available sources; design data models and clustering approaches; and build Python and SQL pipelines that move and organize the data. The engineer will work directly with data and BI analysts to understand source data, reproduce representative structures without using protected CMS data, and validate that the environment supports downstream AI, agentic, and machine-learning proofs of concept. Success requires practical ownership, sound cost control, and a learner-doer mindset.
Requirements
- 5+ years of hands-on data engineering experience, with data engineering as the primary discipline rather than cloud infrastructure alone.
- Greenfield experience standing up and owning a usable data lake and/or data warehouse environment, including architecture, schemas, tables, ingestion, and pipelines.
- Hands-on AWS experience sufficient to help provision and configure a cloud-based data environment and its supporting services.
- Strong Python and SQL development skills for data movement, transformation, and analysis support.
- Practical data modeling and warehouse design experience, including schema design, clustering, and designing for appropriate scale.
- Experience ingesting and normalizing data from flat files, Excel, and/or other varied source formats.
- Working knowledge of Infrastructure as Code or automation using Terraform, OpenTofu, Ansible, or a comparable tool.
- Ability to work independently, control cloud costs, avoid unnecessary overengineering, and learn unfamiliar tools while continuing to deliver.
Nice-to-Haves
- Hands-on Databricks experience.
- AWS QuickSight experience.
- Airflow or comparable data-pipeline orchestration experience.
- Healthcare, government, regulated-data, claims, or CMS-related domain exposure; technical capability remains the priority.
- Experience with data quality and lineage concepts or tools, such as Great Expectations, OpenLineage, Marquez, OpenMetadata, or equivalents.
- Synthetic data generation experience using Faker, SDV, or a comparable tool.
- Experience mentoring analysts or helping Python-focused team members strengthen data engineering and database design practices.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Making Data Warehouses Fast: A Developer’s Story
Highest Paying Tech Companies for Developers
Top Big Data Technologies That You Need to Know
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again