> Markdown version of [/jobs/ext/1186930-data-scientist-ai-data-foundations](https://www.wearedevelopers.com/jobs/ext/1186930-data-scientist-ai-data-foundations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist, AI Data Foundations - **Company:** On behalf of Next Deavor - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $114,000.0 - $175,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Microsoft Azure, Software as a Service, Encodings, Data Discovery, Data Infrastructure, Data Structures, Python (Programming Language), Machine Learning, Neo4j, NumPy, Open Source Technology, Azure Data Lake, Search Technologies, SQL Databases, Data Classification, Large Language Models, Indexer, Pandas, Pyspark, Scikit Learn, HuggingFace, Cosmos DB, Machine Learning Operations, Unsupervised Learning, Databricks - **Published:** July 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=68a0f34bb76150df ## About the Role 4-7 years of experience in data science, ML engineering, or applied data roles, with significant time building data assets consumed by models or applications. Hands-on experience designing and operating vector stores for RAG or semantic search (embedding generation, chunking, indexing, retrieval evaluation). Experience building or operating a feature store (e.g., Databricks Feature Store, Feast, or custom), including offline training and online serving patterns and point-in-time correctness. Experience modeling and building graph data structures and writing graph queries (Neo4j, TigerGraph, Cosmos DB Gremlin, or similar). Strong proficiency in Python (pandas, NumPy, scikit-learn, PySpark) and SQL; comfortable using Databricks notebooks and jobs. Practical experience with embedding models and LLM tooling (Hugging Face, OpenAI/Azure OpenAI APIs, LangChain or similar) in production or near-production contexts. Demonstrated data discovery skills: profiling messy datasets, surfacing patterns, validating findings statistically, and explaining results clearly. Solid grounding in classical ML concepts (supervised vs. unsupervised learning, train/test discipline, leakage, evaluation metrics). Strong written and verbal communication skills for technical and business audiences. Here's What Else Might Help You Out Experience in SaaS or FinTech, especially with lending, deposit, credit, fraud, or KYC/AML data. Familiarity with Databricks-native AI/ML tooling: Databricks Vector Search, Databricks Feature Store, MLflow, Unity Catalog. Experience with open-source vector DBs (pgvector, Pinecone, Weaviate, Chroma, FAISS) and strong opinions on trade-offs. Experience with Microsoft Azure data and AI services (Azure OpenAI, Azure AI Search, ADLS Gen2). Experience evaluating RAG systems end-to-end (recall@k, faithfulness, answer quality, hallucination measurement). Exposure to graph algorithms (community detection, link prediction, centrality) applied to business problems. Bachelor's or Master's in CS, Statistics, Mathematics, Engineering, or related quantitative field, or equivalent experience. ## Description You will design and build the curated data structures that AI and ML applications consume, enabling higher-quality model training and inference. You will partner with model builders, product, risk, and growth stakeholders to surface actionable insights and ship production-ready vector, feature, and graph data assets. This is a Remote role. Here's How You'll Make an Impact on the Team Build and maintain vector stores for RAG, including embedding pipelines, chunking strategies, indexing, and refresh patterns. Own the feature store: design, build, and operate feature definitions, freshness SLAs, lineage, and point-in-time correctness for offline/online use. Design and implement graph data structures to model relationships across applicants, applications, products, lenders, decisions, and outcomes. Lead data discovery: profile lending, deposit, and behavioral datasets to identify trends, segments, anomalies, and model drivers; produce actionable hypotheses for stakeholders. Engineer curated, AI-ready datasets with appropriate quality checks, documentation, and governance for downstream model builders and analysts. Define and run evaluation frameworks for RAG retrieval quality, feature drift, embedding quality, and graph completeness; iterate on metrics. Partner closely with ML engineers and applied scientists to ensure data assets accelerate model development and serving workflows. Champion responsible data use by collaborating with governance, security, and compliance teams to ensure data classification, consent, and regulatory boundaries are respected. Communicate findings via write-ups, notebooks, dashboards, and short presentations for technical and non-technical audiences. ## Related Videos - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)