Data Engineer

ONE STOP COLLECTIBLE CORP
New York, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$200,000.0 - $275,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Amazon S3 Apache HTTP Server Databases Information Engineering Data Governance Identity and Access Management Python (Programming Language) PostgreSQL MariaDB
+9 more
MongoDB Operational Databases Business Intelligence Development Studio Snowflake Jupyter Build Management Vertica Functional Programming Databricks

Job description

You will be Doctronic’s first dedicated data engineer, and you will own the plumbing end to end: how data moves from our production systems into our lakehouse and warehouse, how it gets transformed into trusted, documented tables, and who can access what.

This role serves every team in the company: AI engineering, product, finance, partnerships, and data to name a few.

What You’ll Do

  • Build reliable, monitored CDC pipelines from our production databases (MariaDB, PostgreSQL, MongoDB) into our S3 + Iceberg lake and Snowflake
  • Stand up a transformation layer (e.g. dbt) on Snowflake so core business metrics (visits, bookings, revenue, retention) come from tested, version-controlled models
  • Select and implement an orchestration tool so pipelines and dashboard refreshes run automatically, with alerting when they break
  • Design and enforce the access control model for patient data: row/column-level PHI restrictions, HIPAA Safe Harbor compliance, anonymization pipelines, and account deletion workflows
  • Establish a single governed copy of production data that analytics, finance, and the AI team all read from
  • Support the AI team’s data needs for model training
  • Design and build a best-practice warehouse architecture with clean raw, transformed, and business-ready layers powering our executive dashboards

Requirements

  • 5+ years of data engineering experience, including ownership of production data platforms end to end
  • Strong SQL and Python, with experience building and operating ELT/CDC pipelines (Fivetran, Airbyte, or similar)
  • Hands-on experience with a modern lakehouse/warehouse stack: S3, Apache Iceberg, a catalog layer, and Snowflake or an equivalent warehouse
  • Experience with transformation frameworks (dbt or similar) and orchestration tools (Airflow, Dagster, Glue workflows, or similar)
  • Solid AWS fundamentals: IAM, Lambda, Kinesis, Glue
  • A pragmatic, reliability-first mindset
  • Comfort operating with high autonomy and minimal specs in a flat, engineering-first organization
  • Strong communication skills; you’ll work directly with product, marketing, finance, and AI stakeholders

Nice to Have

  • Experience with HIPAA/PHI data governance, anonymization, or healthcare data
  • Experience with event/behavioral data pipelines (ClickHouse, GTM/server-side tracking, CDPs)
  • Familiarity with ML data workflows: feature pipelines, training datasets, notebook environments (SageMaker, Databricks, Jupyter)
  • Experience with BI tooling (Metabase or similar) and semantic/metrics layers
  • Prior experience as the first or only data engineer at a startup

Benefits & conditions

  • Base salary range: $200,000 to $275,000 annually, depending on experience, plus meaningful equity

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:31 min

Producing candlestick visualization charts inside integrated Jupyter notebooks

Akmal Chaudhri Akmal Chaudhri · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

5:36 min

Building a data analysis stack with Python and Jupyter

Markus Harrer Markus Harrer · World Congress 2021

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all