Lead Data Scientist (Context Engineering)
Rivian · Irvine, CA · 2 wk ago
On-siteAnalyst$154k–$193k/yrFull-time
About the role
Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we constantly challenge what’s possible, never simply accepting what has always been done. We reframe old problems, seek new solutions and operate comfortably in areas that are unknown. Our backgrounds are diverse, but our team shares a love of the outdoors and a desire to protect it for future generations.
Responsibilities
- : Design, develop, and own the semantic layer and business ontology—mapping definitions, relationships, and business logic across Marketing, Sales, and Fulfillment.
- Ensure AI agents have a standardized, enterprise-wide conceptual framework to reason accurately and consistently.
- Cross-Functional Data Modeling: Architect a unified 'Customer 360' schema that harmonizes disparate data streams from Marketing, Sales, and Fulfillment into a single, AI-consumable source of truth.
- RAG & Vector Architecture: Design the Retrieval-Augmented Generation (RAG) infrastructure, determining how unstructured data (e.g., shipping logs, sales notes, brand guidelines) is indexed, chunked, and retrieved for agentic context.
- Identity & State Orchestration: Build the logic for persistent identity resolution and conversation state, ensuring an AI agent can maintain customer context as they move from a sales inquiry to a fulfillment update.
- Agentic Tooling & API Design: Architect the Action Layer—the secure framework of APIs and function calls that allow agents to execute tasks across Salesforce, Braze, and specialized in-house applications.
- AI Data Governance: Establish protocols for data privacy, security, and latency, ensuring agents only access authorized data and respond with sub-second retrieval times.
Qualifications
- Education: Bachelor's degree in Computer Science, Information Systems, Data Engineering, or a related technical field; Master's degree preferred.
- Experience: 8+ years in Data Architecture, Systems Engineering, or Data Engineering, with a proven track record of designing complex, multi-source data ecosystems.
- Semantic Modeling & Ontologies: Proven experience building semantic layers, knowledge graphs, or data ontologies. Ability to translate abstract business rules and cross-departmental relationships into structured data definitions that LLMs can naturally interpret.
- Architectural Vision: A systems-thinking mindset; you can map how a change in a fulfillment status ripples through the data layer to inform a marketing re-engagement agent.
- Databricks Expertise: Deep mastery of the Databricks ecosystem (Unity Catalog, Delta Lake, Vector Search) to manage the end-to-end data lifecycle for AI.
- Technical Mastery: Expert-level command of vector databases and orchestration frameworks (e.g., LangChain, LlamaIndex). Expert-level Python and strong SQL. You understand the plumbing required to connect LLMs to enterprise data.
- MarTech & CRM Literacy: Deep functional knowledge of how data is structured within CDP platforms, CRM systems (e.g., Salesforce), and behavioral event streams to support cross-departmental agent hand-offs.
- Internal Systems Integration: Experience architecting data flows between modern cloud stacks and proprietary in-house applications, ensuring seamless bi-directional communication for AI agents.
- Environment: Comfortable operating in a fast-paced, high-ambiguity environment with strong attention to detail, a builder's mentality, and the ability to make principled architectural decisions with incomplete information.
- Ability to stand, sit, or walk for 8-10 hours per day.