Data Engineer

Tech Quarry (SATO), LLC
United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Microsoft Excel Artificial Intelligence Airflow Amazon Web Services Cloud Computing Cluster Analysis Information Engineering Extract Transform Load (ETL) Data Normalization Data Warehousing Database Design Database Development
+11 more
Python (Programming Language) Machine Learning Operational Databases Ansible SQL Databases Data Lakes Core Data Terraform Data Pipelines Databricks Data Generation

Job description

We are hiring a Data Engineer to build a lightweight, non-production data environment that mirrors the functionality and core data structures of a partner environment at a smaller scale. This person will stand up the AWS-based environment in partnership with IT, configure Databricks and QuickSight, create an initial data lake, and progressively develop warehouse schemas and tables for analytics and AI experimentation. The role will ingest data from flat files, Excel, and other available sources; design data models and clustering approaches; and build Python and SQL pipelines that move and organize the data. The engineer will work directly with data and BI analysts to understand source data, reproduce representative structures without using protected CMS data, and validate that the environment supports downstream AI, agentic, and machine-learning proofs of concept. Success requires practical ownership, sound cost control, and a learner-doer mindset.

Requirements

  • 5+ years of hands-on data engineering experience, with data engineering as the primary discipline rather than cloud infrastructure alone.
  • Greenfield experience standing up and owning a usable data lake and/or data warehouse environment, including architecture, schemas, tables, ingestion, and pipelines.
  • Hands-on AWS experience sufficient to help provision and configure a cloud-based data environment and its supporting services.
  • Strong Python and SQL development skills for data movement, transformation, and analysis support.
  • Practical data modeling and warehouse design experience, including schema design, clustering, and designing for appropriate scale.
  • Experience ingesting and normalizing data from flat files, Excel, and/or other varied source formats.
  • Working knowledge of Infrastructure as Code or automation using Terraform, OpenTofu, Ansible, or a comparable tool.
  • Ability to work independently, control cloud costs, avoid unnecessary overengineering, and learn unfamiliar tools while continuing to deliver.

Nice-to-Haves

  • Hands-on Databricks experience.
  • AWS QuickSight experience.
  • Airflow or comparable data-pipeline orchestration experience.
  • Healthcare, government, regulated-data, claims, or CMS-related domain exposure; technical capability remains the priority.
  • Experience with data quality and lineage concepts or tools, such as Great Expectations, OpenLineage, Marquez, OpenMetadata, or equivalents.
  • Synthetic data generation experience using Faker, SDV, or a comparable tool.
  • Experience mentoring analysts or helping Python-focused team members strengthen data engineering and database design practices.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

Videos

See all

Related articles

See all