> Markdown version of [/jobs/ext/2677164-ce-data-architect](https://www.wearedevelopers.com/jobs/ext/2677164-ce-data-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # CE Data Architect - **Company:** Eli Lilly and Company - **Location:** Indianapolis, IN, United States - **Experience:** Expert - **Salary:** $132,000.0 - $193,600.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Data Analysis, Automation of Tests, Microsoft Azure, Cloud Computing, Information Systems, Continuous Integration, Data Architecture, Information Engineering, Data Governance, Dimensional Modeling, Python (Programming Language), Role-Based Access Control, SQL Databases, Data Streaming, Systems Integration, Tokenization, Policy as Code, Data Classification, Genesys, Git, Pytest, Pyspark, Information Technology, Terraform, Crosswalk, User Identification, Databricks - **Published:** September 2, 2026 - **Apply:** https://www.biospace.com/logon?PipelinedPage=%2Fjob%2F3071897%2Fce-data-architect%3FAction%3DContinueJobApplication%23application-form ## About the Role * Bachelor's degree in Computer Science, Information Systems, Data Engineering, or a related field. * 5+ years of experience in data architecture and/or data engineering, including enterprise-scale delivery. * SQL and hands-on Databricks / Unity Catalog fluency (notebooks, Delta, catalog/schema/grants). * Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization or visas for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT, J-1, H-1B, TN, O-1, E-3, H-1B1, or L-1. Additional Skills / Preferences * Proficiency with Git-based CI/CD workflows (Azure DevOps or equivalent) for versioned data, contract, and policy artifacts. * Demonstrated experience designing lakehouse / medallion architectures (Databricks, Unity Catalog, Delta). * Hands-on data modeling across conceptual, logical, and physical layers, including canonical/converged and dimensional modeling. * Experience with MDM / identity resolution (Reltio or comparable) and cross-system identifier crosswalks. * Working knowledge of data governance, classification, and access control (RBAC/ABAC, row- and column-level security). * Experience in regulated or healthcare data environments; familiarity with HIPAA, PHI handling, and covered-entity constructs. * Data contract frameworks (ODCS or comparable) and policy-as-code (OPA/Rego) awareness. * Tokenization / de-identification patterns (Datavant or comparable). * Pharmacy or patient data flows and consent platforms (Transcend / Cassie). * Ability to translate architecture clearly for legal, compliance, and engineering audiences. * Cloud/data architecture certifications (AWS, Databricks). * Python / PySpark and transformation-as-code (dbt) for reference pipelines, with a test framework (pytest, DLT / Great Expectations). * Infrastructure-as-code (Terraform) for provisioning Unity Catalog, grants, and the CE isolation boundary. * Experience architecting or building AI Skills and agents and operationalizing them through CI/CD. ## Description The CE Data Architect owns the end-to-end data architecture for the CE Pharmacy platform within the covered-entity boundary: the canonical data models, the medallion/lakehouse structure, the identity crosswalk, PHI classification, data contracts, and the isolation controls that keep identified PHI from crossing to parent-side systems without tokenization. This is a governance-first, evidence-based role that translates regulatory and privacy constraints into durable, load-bearing architecture - partnering closely with stewardship, policy-as-code, and engineering to ensure the platform is compliant, auditable, and built to scale., * Own the canonical CE data architecture - conceptual, logical, and physical models across the Databricks / Unity Catalog medallion (bronze * silver * gold), scoped entirely within the CE trust boundary. * Model the CE identity and crosswalk layer spanning Reltio MDM, Auth0/Passport, Datavant tokens, Scriptly, Genesys, Prescryptive, and Transcend - defining authoritative-anchor logic and safeguarding merge/split integrity. * Author and steward data contracts (ODCS) for CE data products, defining schema, quality, freshness/SLA, and lineage expectations between producers and consumers. * Codify PHI data classification across the 300-field registry and ensure row/column-level security, masking, and tokenization boundaries are enforced architecturally - not merely documented. * Design the CE isolation model that prevents identified PHI from crossing to parent-side systems without Datavant tokenization; partner with the Policy-as-Code Engineer to render architectural intent as OPA/Rego. * Establish reference architecture and patterns so Analytics Engineers and Data Engineers build CE data products consistently and to standard. * Define lineage, cataloging, and metadata strategy for CE, ensuring end-to-end traceability that supports audit readiness and incident response. * Provide architectural review and gatekeeping for new CE data flows, integrations, and AI/agent consumption paths. * Partner across governance and platform - HIPAA governance, Privacy, Legal/DLO, and platform owners (PPH) - to keep the architecture compliant and defensible. * Enable Skill and agent development - architect the data contracts, reference patterns, and guardrails that let Applied-AI / Agent Engineers build CE Skills and agents against governed data, so PHI classification and consent travel with the data into agentic consumption. * Supply CI/CD patterns for data, policy, and agent artifacts - define and standardize the pipelines (Azure DevOps, Git-based promotion across dev * test * prod) that version and ship models, ODCS contracts, OPA/Rego policies, Skills, and agents within the CE boundary. ## Related Videos - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [pytest: Simple, rapid and fun testing with Python](https://www.wearedevelopers.com/videos/213-pytest-simple-rapid-and-fun-testing-with-python) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)