Lead Data Scientist (Context Engineering)

Rivian
Irvine, CA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Data Analysis Business Logic Systems Engineering Information Systems Data Architecture Information Engineering Data Governance Data Infrastructure Graph Database Python (Programming Language)
+17 more
Salesforce.Com Search Technologies Data Streaming Systems Integration Unstructured Data Enterprise Data Management Large Language Models Generative AI Data Layers Data Lakes Kubernetes Information Technology Low Latency Virtual Agents Api Design User Identification Databricks

Job description

Rivian’s AI team sits within the Analytics function of the Customer organization-the team that spans Marketing, Sales, and Market Intelligence-and is building the semantic and data infrastructure that makes intelligent agents possible at scale. As a Staff/Lead Data Scientist focused on Context Engineering, you will design the knowledge layer that AI agents reason from-owning the ontologies, semantic models, identity resolution logic, and retrieval architecture that unify Rivian’s Marketing, Sales, and Fulfillment data into a coherent, AI-consumable ecosystem. This is a foundational systems role: the agents are only as good as the context you build. You’ll work at the intersection of data architecture, knowledge engineering, and applied AI, partnering closely with the Agentic Solutions team, Data Platform, MarTech, and the broader Customer org Analytics team. * Customer Domain Ontology: Design, develop, and own the semantic layer and business ontology-mapping definitions, relationships, and business logic across Marketing, Sales, and Fulfillment. Ensure AI agents have a standardized, enterprise-wide conceptual framework to reason accurately and consistently.

  • Cross-Functional Data Modeling: Architect a unified ‘Customer 360’ schema that harmonizes disparate data streams from Marketing, Sales, and Fulfillment into a single, AI-consumable source of truth.
  • RAG & Vector Architecture: Design the Retrieval-Augmented Generation (RAG) infrastructure, determining how unstructured data (e.g., shipping logs, sales notes, brand guidelines) is indexed, chunked, and retrieved for agentic context.
  • Identity & State Orchestration: Build the logic for persistent identity resolution and conversation state, ensuring an AI agent can maintain customer context as they move from a sales inquiry to a fulfillment update.
  • Agentic Tooling & API Design: Architect the Action Layer-the secure framework of APIs and function calls that allow agents to execute tasks across Salesforce, Braze, and specialized in-house applications.
  • AI Data Governance: Establish protocols for data privacy, security, and latency, ensuring agents only access authorized data and respond with sub-second retrieval times.

Requirements

Education: Bachelor’s degree in Computer Science, Information Systems, Data Engineering, or a related technical field; Master’s degree preferred.

  • Experience: 8+ years in Data Architecture, Systems Engineering, or Data Engineering, with a proven track record of designing complex, multi-source data ecosystems.
  • Semantic Modeling & Ontologies: Proven experience building semantic layers, knowledge graphs, or data ontologies. Ability to translate abstract business rules and cross-departmental relationships into structured data definitions that LLMs can naturally interpret.
  • Architectural Vision: A systems-thinking mindset; you can map how a change in a fulfillment status ripples through the data layer to inform a marketing re-engagement agent.
  • Databricks Expertise: Deep mastery of the Databricks ecosystem (Unity Catalog, Delta Lake, Vector Search) to manage the end-to-end data lifecycle for AI.
  • Technical Mastery: Expert-level command of vector databases and orchestration frameworks (e.g., LangChain, LlamaIndex). Expert-level Python and strong SQL. You understand the plumbing required to connect LLMs to enterprise data.
  • MarTech & CRM Literacy: Deep functional knowledge of how data is structured within CDP platforms, CRM systems (e.g., Salesforce), and behavioral event streams to support cross-departmental agent hand-offs.
  • Internal Systems Integration: Experience architecting data flows between modern cloud stacks and proprietary in-house applications, ensuring seamless bi-directional communication for AI agents.
  • Environment: Comfortable operating in a fast-paced, high-ambiguity environment with strong attention to detail, a builder’s mentality, and the ability to make principled architectural decisions with incomplete information.
  • Ability to stand, sit, or walk for 8-10 hours per day.
  • Required to communicate using phone and/or e-mail.
  • Ability to view, read, and interpret documents.
  • Ability to perform all duties in an office environment that may contain ambient noise and temperature fluctuations.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:27 min

Managing traffic and tracking costs with Databricks Unity Catalog

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

3:10 min

Understanding the core concepts of API design

Alen Pokos · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:50 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

1:50 min

Introduction to the speaker and data science background

Bas Geerdink · LIVE

Videos

See all

Related articles

See all