> Markdown version of [/jobs/ext/2954139-data-engineer](https://www.wearedevelopers.com/jobs/ext/2954139-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Eli Lilly and Company - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $157,500.0 - $231,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Advanced Message Queuing Protocol, Automation of Tests, Microsoft Azure, Computer Clusters, Databases, Continuous Integration, Data Transmissions, Information Engineering, Extract Transform Load (ETL), Data Systems, Distributed Computing Environment, Experimental Data, Python (Programming Language), PostgreSQL, Linked Data, Machine Learning, Message Queuing Telemetry Transport (MQTT), Operational Databases, Standard Sql, Software Systems, Data Streaming, Management of Software Versions, Parquet, Apache Spark, Git, Information Technology, Data Analytics, Dask, Apache Kafka, Data Management, Data Pipelines - **Published:** September 17, 2026 - **Apply:** https://www.biospace.com/logon?PipelinedPage=%2Fjob%2F3074908%2Fdata-engineer%3FAction%3DContinueJobApplication%23application-form ## About the Role * Strong Python, or equivalent experience building data-intensive software systems. * Strong SQL and data modeling experience including designing schemas that hold up as scientific data grows and diversifies, with expert knowledge of Postgres or a comparable enterprise database. * Distributed data processing (Spark, Ray, or Dask) and pipeline orchestration (Airflow or Dagster) at scale. * Experience with cloud platforms - AWS and Azure preferred - and with high-performance and object storage feeding large-scale compute environments. * A track record of building data systems that other people depend on, and of taking responsibility for them when they broke. * Strong testing practices and test automation, with solid CI/CD and Git fundamentals. * Adaptability and a collaborative mindset, with the ability to translate complex scientific questions into data solutions that accelerate experimentation and decision-making. * Experience streaming and event-driven integration (Kafka, MQTT, or AMQP), including instrument and laboratory data capture. * Cheminformatics or scientific data experience - compound registration, structure notation (SMILES, InChI, HELM), RDKit, multi-omics, assay, or sequencing data - is a strong plus. * Prior experience across the following: data modeling, ETL/ELT at scale, ontology development, semantic graph construction and linked data, or relational schema design. * Experience standing up, migrating, or consolidating databases and data platforms, including production cutover of systems in active use. Your Basic Qualifications * Bachelor's degree in Computer Science, Data Science, Engineering, Mathematics, or a related technical field. * 5+ years of data engineering experience building and operating production data systems. ## Description As a Data Engineer, you will build and maintain the data platforms that power AI-driven research and discovery. You will develop scalable pipelines that ingest, transform, and deliver chemical, biological, and experimental data for machine learning and scientific workflows. Partnering with AI Scientists, AI Engineers, and laboratory researchers, you will ensure that data is accurate, traceable, and accessible at scale. Your work will provide the trusted data foundation behind next-generation AI models and experiments. How You'll Succeed * Engineer datasets in large language environment for model training specifically efficient formats and storage layout (Parquet, Zarr, Arrow) and delivery fast enough that GPU clusters are never left waiting on data. * Design, develop, and maintain scalable and efficient data pipelines to support data analytics, reporting, and machine learning initiatives. * Ensure seamless data flow between systems and applications, optimizing data transfer and transformation processes for performance and scalability. * Build the ingestion path from the automated lab, so experimental results reach the models in hours rather than weeks, closing the loop between what a model proposes and what the next model learns from. * Own the correctness of what models train on completeness, sound joins across experimental sources, and validation that catches a bad dataset before it reaches a training run rather than after. * Build dataset versioning, lineage, and reproducibility into the platform, so any model can be traced to the exact data it was trained on months or years later. * Work with the laboratory, instrument, and external teams producing the data so that a change upstream does not quietly corrupt a training run downstream. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)