> Markdown version of [/jobs/ext/2281203-agentic-ai-data-engineer-cmc-data-integration](https://www.wearedevelopers.com/jobs/ext/2281203-agentic-ai-data-engineer-cmc-data-integration). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Agentic AI Data Engineer - CMC Data Integration - **Company:** Eli Lilly and Company - **Location:** Indianapolis, IN, United States - **Salary:** $65,250.0 - $169,400.0 - **Contract:** Permanent contract - **Skills:** Microsoft Excel, Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Audit Trail, Microsoft Azure, Computer Engineering, Information Engineering, Data Infrastructure, Data Integration, Data Integrity, Extract Transform Load (ETL), Data Normalization, Relational Databases, Github, JSON, Python (Programming Language), Laboratory Information Management Systems, Machine Learning, Cloud Services, SQL Databases, Extensible Markup Language (XML), Data Logging, Data Ingestion, Azure Data Factory, Large Language Models, Build Management, Information Technology, Production Code, Machine Learning Operations, Virtual Agents, Data Pipelines, GXP, Databricks - **Published:** August 28, 2026 - **Apply:** https://www.biospace.com/logon?PipelinedPage=%2Fjob%2F3070730%2Fagentic-ai-data-engineer-cmc-data-integration%3FAction%3DContinueJobApplication%23application-form ## About the Role * MS in Computer Science, Computer Engineering, Data Engineering, or related technical field with 1-2 years of relevant experience; OR * BS in Computer Science or Computer Engineering with 3-5 years of hands-on data engineering experience. * Proficiency in Python and SQL; ability to write, review, and own production-quality code. * Demonstrated experience building ETL/ELT pipelines from unstructured or semi-structured sources (PDFs, Excel, JSON, XML). * Hands-on experience building LLM-powered applications: retrieval-augmented generation, tool-calling, multi-step orchestration, or equivalent agentic patterns. * Hands-on experience with cloud data platforms: Azure (Data Factory, Databricks, Fabric) or AWS (S3, Glue, Lambda, Redshift). * Solid understanding of relational data modeling, schema design, and data normalization principles. * Familiarity with data orchestration tools (Airflow, Azure Data Factory, Prefect, or similar). * Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization or visas for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT, J-1, H-1B, TN, O-1, E-3, H-1B1, or L-1. Additional Preferences: * Working knowledge of 21 CFR Part 11, ALCOA+, and GxP data integrity principles, or clear demonstrated ability to apply similar audit/compliance frameworks. * Experience integrating data from LIMS, ELN, SDMS, or CDS systems (Benchling, LabVantage, OpenLABS, or equivalent). * Familiarity with pharmaceutical CMC data types: analytical results, batch records, stability studies, specifications. * Experience with data mesh architecture or data product ownership models. * Knowledge of MLOps practices and preparing data for AI/ML model training in regulated environments. * Exposure to regulatory submission data formats (eCTD, CTD, CDISC SEND/SDTM). * Experience with CI/CD pipelines (GitHub Actions, Azure DevOps) applied to data engineering workloads. ## Description You will work with a team of engineers and data scientists. You will have the autonomy to own your components end-to-end. If you want hands-on experience at the intersection of pharmaceutical science and modern agentic AI data engineering - agentic pipelines, document AI, GxP-compliant data infrastructure - this is the role to build that foundation., Agentic Pipeline Components: * Implement individual agent components (e.g., document extraction agent, schema mapping agent, validation agent) within the established orchestration framework (LangGraph, LlamaIndex, or equivalent) * Write tool-calling logic, handle failure modes, and ensure each agent component is testable and observable with instrumented logging of inputs, outputs, and intermediate decisions * Iterate on agent behavior based on real data performance; work with the senior engineer to identify and resolve failure patterns * Participate in validation and qualification activities for AI-assisted workflows, supporting documentation that demonstrates computational tools reflect scientific intent Human-in-the-Loop (HITL) Workflow Implementation: * Build review queues and flagging logic that surface low-confidence or out-of-specification extractions to scientific reviewers for approval before data is loaded * Implement routing logic that captures reviewer decisions, logs outcomes with full audit trail, and reintegrates approved data into the pipeline per 21 CFR Part 11 electronic records requirements * Tune flagging thresholds based on feedback from scientific owners; maintain and improve HITL logic as new data sources are onboarded Data Ingestion & Pipeline Engineering: * Design and build AI-assisted ingestion pipelines that extract and structure the data from unstructured CDMO/CRO data sources: PDFs (Certificates of Analysis, batch records), Excel files, and vendor portal exports * Implement validation, reconciliation, and exception-handling logic to ensure data completeness and integrity before loading * Build monitoring and alerting for pipeline health, data quality, and ingestion failures * Design a data quality framework with automated checks, rejection handling, and audit trail logging. * Develop reusable pipeline templates and schema documentation that reduce onboarding time for new CDMO partners ## Related Videos - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)