> Markdown version of [/jobs/ext/2727957-data-engineer](https://www.wearedevelopers.com/jobs/ext/2727957-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Cohort AI Inc - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Adobe InDesign, Artificial Intelligence, Airflow, Amazon Web Services, Microsoft Azure, Bash Shell, Big Data, BigQuery, Health Informatics, Clinical Data Repository, Cloud Computing, Software Quality, Code Review, Continuous Integration, Information Engineering, Data Infrastructure, Data Integration, Extract Transform Load (ETL), Data Systems, Data Warehousing, Database Queries, Linux, Distributed Computing Environment, Distributed Data Store, Distributed Systems, Python (Programming Language), Linux Commands, Operational Databases, Shell Script, SQL Databases, Data Processing, Fast Healthcare Interoperability Resources, Large Language Models, Snowflake, Apache Spark, Generative AI, Git, Pyspark, Health Level Seven International, Software Version Control, Data Pipelines, Amazon Redshift, Databricks - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/senior-data-engineer-cohort-ai-inc-9932598 ## About the Role * 5+ years of professional Data Engineering experience. * Strong SQL skills and experience working with large datasets. * Strong hands-on experience with Python and PySpark. * Experience working with Apache Spark and distributed data processing. * Hands-on experience with modern data platforms such as Databricks, Snowflake, BigQuery, Redshift, or similar technologies. * Experience designing and operating production-grade ETL/ELT pipelines. * Strong understanding of data modeling, data warehousing, and distributed data systems. * Solid Linux command-line experience, including bash and shell scripting. * Experience implementing data quality, monitoring, and alerting frameworks. * Experience with Git and CI/CD. * Strong problem-solving skills and the ability to independently investigate and resolve complex technical issues. * Strong communication skills and the ability to collaborate effectively with cross-functional teams. Nice to Have * Experience working with healthcare data and standards such as OMOP, FHIR, HL7, ICD, CPT, claims, EHR/EMR, or related datasets. * Experience with AWS, GCP, or Azure. * Experience working with large-scale distributed systems. * Familiarity with Airflow, Dagster, Prefect, or similar workflow orchestration tools. * Exposure to Generative AI, LLMs, or AI-enabled data applications. * Experience working in healthcare, life sciences, health technology, or a data-intensive environment ## Description * Design, build, and maintain scalable, reliable data pipelines. * Develop high-performance ETL/ELT workflows using SQL, Python, PySpark, and Apache Spark. * Work with complex healthcare datasets, including clinical and claims data. * Build ingestion and transformation workflows that are reliable, maintainable, and scalable. * Develop and improve data quality, monitoring, alerting, and observability solutions. * Troubleshoot and resolve production data pipeline issues using Linux and shell-based tools. * Optimize data processing performance and cloud infrastructure costs. * Apply strong engineering practices around code quality, testing, version control, and CI/CD. * Contribute to data modeling, data warehousing, and distributed data architecture decisions. * Work closely with Data Science, Clinical Informatics, Product, Infrastructure, and Commercial teams to translate requirements into effective data solutions. * Support customer onboarding and complex data integration initiatives. * Participate in design and code reviews and contribute to improving engineering practices. * Help identify opportunities to improve the scalability, reliability, and efficiency of our data platform. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)