> Markdown version of [/jobs/ext/212770-data-scientist-engineer](https://www.wearedevelopers.com/jobs/ext/212770-data-scientist-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist / Engineer - **Company:** Shift - **Location:** Boston, MA, United States - **Salary:** $80,000.0 - $110,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Microsoft Azure, Data Architecture, Machine Learning, Open Source Technology, Operational Databases, Systems Integration, Management of Software Versions, Large Language Models, Prompt Engineering, Apache Spark, Model Validation, Data Lakes, Machine Learning Operations, Virtual Agents, GPT, Data Pipelines, Databricks - **Published:** May 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=af9c79e11a78c858 ## About the Role We are looking for candidates with diverse skills to help us build excellent technology solutions for our clients and be proficient in the following skills: * Expert proficiency in production-level object-oriented programming (OOP) for building scalable and reliable systems. * Proven hands-on experience with Large Language Models (LLMs) and generative AI techniques (including RAG, embeddings, prompt engineering, and model tuning), leveraging frameworks such as OpenAI/Anthropic or open-source variants. * Solid foundation in ML fundamentals with practical experience in the full machine learning lifecycle, including model evaluation, monitoring, versioning, and deployment in production environments. * Experience designing and implementing robust data pipelines for document, OCR, and multi-modal data workflows. * Agentic Frameworks: Experience with integrating frameworks like LangChain/LangGraph, OpenAI Agent SDK, CrewAI, A2A/MCP with Databricks or Azure-hosted models (e.g., DBRX, OpenAI GPT-5.3, Anthropic Claude, Google Gemini). * Demonstrated ability to effectively engage with clients, translate complex business needs into clear, actionable technical solutions, and manage stakeholder expectations. Highly Desired Skills * Databricks Ecosystem: Deep expertise in Mosaic AI (formerly MosaicML), Unity Catalog, and Delta Lake. Good understanding of Spark data architecture is a plus. * Experience using MLflow for the full lifecycle: from experiment tracking and prompt engineering in the AI Playground to model evaluation. ## Description * Your role will be to actively contribute to the US- Insurance roadmap and clients, and working on various data types such as structured data, free text, documents and images. * Build and productionize data pipelines (structured, text, documents, images) optimized for LLMs and multi-modal models. * Design, develop and deploy LLM-based solutions (RAG, embeddings, instruction tuning) for subrogation, claims handling, document understanding, and related use cases. * Experiment with the latest in Agentic AI technologies (Langchain/Langgraph, OpenAI Agent SDK, MCP, A2A) and develop MVP for the next generation of autonomous subrogation solutions. * Develop "Chain-of-Thought" and "ReAct" prompting strategies to ensure the agent can justify its liability percentages based on the Comparative Negligence laws of different jurisdictions. * Create custom "tools" for the agent, allowing it to query internal databases, call external weather APIs, or calculate impact force based on telemetry data. * Establish rigorous evaluation frameworks (LLM-as-a-judge) to ensure the agent's decisions are unbiased, legally sound, and explainable. * Ensure responsible-AI practices: privacy, hallucination mitigation, explainability and compliance. * Lead client workshops, present prototypes, gather feedback and help define roadmap priorities. ## Related Videos - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)