> Markdown version of [/jobs/ext/212976-data-scientist](https://www.wearedevelopers.com/jobs/ext/212976-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** The Lilly Company - **Location:** Indianapolis, IN, United States - **Salary:** $66,000.0 - $165,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Amazon Web Services, Amazon S3, Data Integrity, Data Mining, Data Visualization, R (Programming Language), Statistical Hypothesis Testing, Python (Programming Language), Machine Learning, Regression Analysis, Natural Language Processing, Cloud Services, Tensorflow, SAS (Software), Large Language Models, Prompt Engineering, Generative AI, Core Data, Information Technology, Machine Learning Operations, Api Design, Document Classification, GXP, Databricks - **Published:** May 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=ffa8b913c8fc8798 ## About the Role * Bachelor's degree in Data Science, Statistics, Computer Science, Mathematics, or a related quantitative field * 1 years of professional data science experience in Python, R and core data science libraries * Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization now or in the future for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT, J-1, H-1B, TN, O-1, E-3, H-1B1, or L-1. Additional Skills & Preferences * Experience with machine learning frameworks and model deployment patterns * Academic Background in Data Science * Hands-on experience with NLP techniques and/or generative AI - LLM APIs (OpenAI, Anthropic), RAG architectures, vector databases, prompt engineering * Familiarity with cloud data platforms - AWS (SageMaker, Lambda, S3), Databricks, or similar * Knowledge of statistical methods - hypothesis testing, experimental design, Bayesian methods, regression analysis * Experience with SAS programming * Strong communication skills - ability to present technical findings to non-technical audiences and translate business questions into analytical frameworks * Collaborative mindset and experience working with cross-functional teams including engineers, product owners, and business partners ## Description As a Data Scientist, you will work across the portfolio to identify opportunities where data science and AI can create measurable business value. You'll analyze complex datasets - regulatory documents, clinical trial data, submission timelines, operational metrics - to uncover patterns, build predictive models, and develop AI-powered solutions. You'll design and evaluate machine learning models, build NLP and generative AI applications for regulatory and scientific content, and collaborate closely with full stack engineers to move your work from prototype to production. You operate in a regulated, GxP environment where data integrity, reproducibility, and validation are not optional. How You'll Succeed * Partnering with business stakeholders across GRA, GSC, and GSS to understand their workflows, identify high-impact problems, and frame them as data science opportunities - translating business questions into analytical approaches. * Developing and deploying machine learning models - classification, regression, clustering, time-series forecasting - to solve problems such as submission timeline prediction, document classification, regulatory risk scoring, and resource optimization. * Building and evaluating NLP and generative AI solutions - leveraging LLMs, RAG architectures, text extraction, entity recognition, and document summarization to automate regulatory authoring, scientific literature analysis, and content generation workflows. * Designing and executing experiments to evaluate model performance - using rigorous statistical methods, A/B testing, and evaluation frameworks (including RAGAS for RAG systems) to ensure solutions meet quality and accuracy thresholds before deployment. * Designing and building AI agents and agentic workflows - creating multi-step, tool-using systems that can autonomously execute complex tasks such as regulatory document drafting, data extraction and transformation, and cross-system orchestration - moving beyond single-prompt interactions to production-grade agent architectures that operate reliably in a validated environment. * Collaborating with full stack engineers and platform teams to productionize models - building APIs, integrating into existing applications, deploying on AWS infrastructure (Lambda, EKS, SageMaker, Databricks), and monitoring model performance in production. * Communicating findings and recommendations to both technical and non-technical audiences - using data visualization, storytelling, and clear business-impact framing to ensure your work drives actual decisions. * Staying current with emerging techniques in machine learning, generative AI, and data science - evaluating new tools, frameworks, and approaches for applicability to the GRA/GSC/GSS portfolio and sharing knowledge with the broader team., * Will work a hybrid schedule in Indianapolis, IN * Remote employees will not be considered * Lilly will not provide support for or sponsor work authorization now or in the future for this role Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response. ## Related Videos - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Building Multi-Tenant ASP.NET Core Applications: Best Practices and Real-World Solutions](https://www.wearedevelopers.com/videos/1552-building-multi-tenant-asp-net-core-applications-best-practices-and-real-world-solutions) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)