> Markdown version of [/jobs/ext/2706573-data-engineer-bioinformatics](https://www.wearedevelopers.com/jobs/ext/2706573-data-engineer-bioinformatics). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer, Bioinformatics - **Company:** Lila Sciences, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $144,000.0 - $240,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Computational Biology, Information Engineering, Extract Transform Load (ETL), Data Transformation, Relational Databases, Database Queries, Python (Programming Language), Laboratory Information Management Systems, PostgreSQL, NumPy, Data Streaming, Workflow Management Systems, Parquet, Pandas, Production Code, Apache Kafka, Data Pipelines - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-data-engineer-bioinformatics-cheminformatics-materials-lila-sciences-9838266 ## About the Role * 2-6 years of experience in data engineering, bioinformatics, cheminformatics, or computational science. * Strong Python skills, including typed, tested, production-quality code. * Strong SQL skills, especially with Postgres or similar relational databases. * Experience building ETL pipelines, data models, and reusable data transformations. * Data science foundation, including statistics and pandas, NumPy, or similar tools. * Experience translating noisy scientific measurements into accurate, validated datasets. * Workflow orchestration experience, ideally Flyte, Airflow, Prefect, Dagster, or Nextflow. * Active use of AI coding tools in day-to-day engineering work. Bonus Points For * Experience with columnar or lakehouse stacks such as Parquet, Iceberg, DuckDB, Polars, or Ibis. * Familiarity with event-driven pipelines such as NATS or Kafka. * Exposure to lab instrument data formats, LIMS, or ELN systems. * Familiarity with life sciences assays, sequencing, imaging, or flow cytometry. * Familiarity with materials or chemistry methods such as XRD, XRF, SEM, TGA, or DSC. * Experience with curve fitting, peak detection, or unit and dimensional analysis. ## Description Lila's mission is to accelerate scientific discovery with AI, and that depends on trustworthy scientific data. As a Data Engineer, you'll build ETL pipelines and data models for Lila's scientific data platform, working at the intersection of data engineering, computational biology, chemistry, and materials science. You'll partner with AI researchers and experimentalists to turn raw lab instrument outputs into validated, analysis-ready datasets. The core challenge is data modeling: transforming messy, per-instrument measurements into clean, well-typed data that is efficient to query, reliable to use, and ready for downstream analysis. You'll also build domain-specific analysis functions and reusable data pipelines that help scientists and AI researchers move faster without re-deriving bespoke solutions. What You'll Be Building * Design pipelines that turn raw lab output into analysis-ready scientific data. * Model heterogeneous data from bio, chemistry, and materials instruments. * Build validation checks, schema-evolution gates, and data quality workflows. * Develop reusable analysis functions for scientific and AI research workflows. * Improve automation and observability across instrument-to-result data flows. * Build canonical datasets that scientists and AI researchers can trust. * Use AI coding tools to accelerate pipeline development and team velocity. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)