Data Engineer

Cohort AI Inc
United States
2 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Adobe InDesign Artificial Intelligence Airflow Amazon Web Services Microsoft Azure Bash Shell Big Data BigQuery Health Informatics Clinical Data Repository Cloud Computing Software Quality
+31 more
Code Review Continuous Integration Information Engineering Data Infrastructure Data Integration Extract Transform Load (ETL) Data Systems Data Warehousing Database Queries Linux Distributed Computing Environment Distributed Data Store Distributed Systems Python (Programming Language) Linux Commands Operational Databases Shell Script SQL Databases Data Processing Fast Healthcare Interoperability Resources Large Language Models Snowflake Apache Spark Generative AI Git Pyspark Health Level Seven International Software Version Control Data Pipelines Amazon Redshift Databricks

Job description

  • Design, build, and maintain scalable, reliable data pipelines.
  • Develop high-performance ETL/ELT workflows using SQL, Python, PySpark, and Apache Spark.
  • Work with complex healthcare datasets, including clinical and claims data.
  • Build ingestion and transformation workflows that are reliable, maintainable, and scalable.
  • Develop and improve data quality, monitoring, alerting, and observability solutions.
  • Troubleshoot and resolve production data pipeline issues using Linux and shell-based tools.
  • Optimize data processing performance and cloud infrastructure costs.
  • Apply strong engineering practices around code quality, testing, version control, and CI/CD.
  • Contribute to data modeling, data warehousing, and distributed data architecture decisions.
  • Work closely with Data Science, Clinical Informatics, Product, Infrastructure, and Commercial teams to translate requirements into effective data solutions.
  • Support customer onboarding and complex data integration initiatives.
  • Participate in design and code reviews and contribute to improving engineering practices.
  • Help identify opportunities to improve the scalability, reliability, and efficiency of our data platform.

Requirements

  • 5+ years of professional Data Engineering experience.
  • Strong SQL skills and experience working with large datasets.
  • Strong hands-on experience with Python and PySpark.
  • Experience working with Apache Spark and distributed data processing.
  • Hands-on experience with modern data platforms such as Databricks, Snowflake, BigQuery, Redshift, or similar technologies.
  • Experience designing and operating production-grade ETL/ELT pipelines.
  • Strong understanding of data modeling, data warehousing, and distributed data systems.
  • Solid Linux command-line experience, including bash and shell scripting.
  • Experience implementing data quality, monitoring, and alerting frameworks.
  • Experience with Git and CI/CD.
  • Strong problem-solving skills and the ability to independently investigate and resolve complex technical issues.
  • Strong communication skills and the ability to collaborate effectively with cross-functional teams.

Nice to Have

  • Experience working with healthcare data and standards such as OMOP, FHIR, HL7, ICD, CPT, claims, EHR/EMR, or related datasets.
  • Experience with AWS, GCP, or Azure.
  • Experience working with large-scale distributed systems.
  • Familiarity with Airflow, Dagster, Prefect, or similar workflow orchestration tools.
  • Exposure to Generative AI, LLMs, or AI-enabled data applications.
  • Experience working in healthcare, life sciences, health technology, or a data-intensive environment

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all