Data Engineer (NYC Hybrid)

Empassion Health, Inc.
New York, United States
2 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Airflow Amazon Web Services Unit Testing Microsoft Azure Big Data BigQuery Cloud Computing Cloud Database Cloud Storage Directed Acyclic Graph (Directed Graphs) Information Engineering
+8 more
Data Integrity Database Queries Python (Programming Language) SQL Databases Data Streaming Delivery Pipeline Looker Analytics Software Version Control

Job description

  • Partner with teams across the business and external partners to understand data needs and deliver reliable pipelines and models that solve real problems.
  • Build and maintain scalable ingestion and egress pipelines in Airflow and dbt Cloud, ensuring high quality, automated data flows across cloud environments.
  • Implement unit tests and monitoring to guarantee data integrity and reproducibility.
  • Model and transform and structure healthcare datasets into usable formats that power data science models, and other reporting marts
  • Enhance and scale data models with SQL and dbt, ensuring precision and adaptability for new partnerships.
  • Write Python code for Apache Airflow DAGs, components and utilities that orchestrate and monitor data workflows. Build complex pipelines that enable flexible scheduling, conditional logic, and smooth integration across multiple data sources.

Requirements

  • 2+ years in data engineering or analytics engineering with proven ability to build pipelines and scalable workflows.
  • Strong SQL skills for querying large, complex datasets.
  • Proficiency in Python for data engineering tasks (transformations, APIs, automation).
  • Experience with cloud data warehouses and storage (GCP preferred: BigQuery, Cloud Storage, Composer; AWS/Azure equivalents acceptable).
  • Hands on experience with dbt or similar data modeling tools.
  • Comfort working in collaborative dev/staging/prod environments, partnering with Product and Tech to safely test, launch, and anticipate the impact of new changes.
  • Curiosity about operational workflows and a drive to partner with non-technical teams, ensuring data and reporting align with how the business actually runs. You’re not just a spec-taker, you’re part of the solution.
  • A proactive, problem-solving mindset and ability to thrive in fast-paced, iterative environments.
  • Strong communication skills to collaborate with analysts, engineers, and business stakeholders. Nice to have:
  • Knowledge of healthcare data (claims, ADT feeds, eligibility files).
  • Familiarity with Git/GitHub for version control.
  • Early-stage startup experience (seed/Series A), especially mission-driven ones.
  • Experience building semantic layers and data models in Looker (LookML). Ready to Make a Difference? If you’re driven by data, healthcare, and impact, apply and let’s talk!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon ¡ World Congress 2026 Europe

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy ¡ LIVE

2:30 min

Leveraging BigQuery ML for scalable SQL-based segmentation experiments

Julian Joseph ¡ LIVE

3:30 min

Approaching data problems with an engineering and strategy mindset

Becky Gandillon ¡ LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy ¡ World Congress 2024

2:10 min

Why organizations combine big data and machine learning

Ayon Roy ¡ LIVE

Videos

See all

Related articles

See all