> Markdown version of [/jobs/ext/1670388-data-scientist](https://www.wearedevelopers.com/jobs/ext/1670388-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** AnswersNow Inc - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $144,000.0 - $168,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Automated Storage and Retrieval Systems, Hospital Information Systems, Data Infrastructure, Python (Programming Language), Azure Machine Learning, Data Processing, Delivery Pipeline, Large Language Models, Prompt Engineering, Model Validation, Information Technology - **Published:** July 7, 2026 - **Apply:** https://www.builtincolorado.com/auth/login?destination=/job/data-scientist/10067707 ## About the Role * 4+ years of experience in applied data science, ML engineering, or AI engineering in a production environment * Deep understanding of RAG architectures: retrieval systems, embedding models, vector databases (Pinecone, Weaviate, pgvector, or similar), chunking strategies, and context assembly * Experience designing and running evaluation frameworks for AI systems - you've thought hard about how to measure quality in domains where ground truth is ambiguous * Strong Python skills; experience with LLM orchestration frameworks (LangChain, LlamaIndex, or similar) * Clinical NLP experience or healthcare AI background is strongly preferred - you understand why clinical data is different from general text and what that means for AI system design * You think like an engineer and a scientist: you build systems that can be measured, iterated on, and trusted - not black boxes * Strong written communication: you can explain RAG pipeline design to a clinician and explain clinical requirements to an engineer * Genuine interest in the clinical domain - you want to understand Applied Behavior Analysis well enough to build AI that actually helps BCBAs do their jobs Nice to have: * Experience with Amazon Bedrock, SageMaker, or AWS AI/ML services * Familiarity with HIPAA-compliant data handling for AI training and inference pipelines * Background in clinical NLP, behavioral health informatics, or ABA/autism research * Experience with fine-tuning or RLHF - even if this role doesn't require it, understanding the tradeoffs informs better RAG design * Exposure to LLM-as-judge evaluation patterns or multi-model evaluation pipelines ## Description As our Data Scientist, you will optimize the smart systems that pull real-time clinical context and turn it into safe, accurate, and highly relevant recommendations. By bridging the gap between cutting-edge AI capabilities and deep clinical expertise, you ensure our models are deeply rooted in real-world care and held to the highest quality standards. Your work ensures that our digital systems are a reliable, trusted partner for our clinical teams, allowing us to safely scale our platform and deliver life-changing autism therapy to families nationwide., * Architect and continuously improve the RAG pipeline that retrieves client-specific clinical context - session notes, treatment plan goals, historical performance data - and injects it into inference-time prompts * Design the retrieval layer: chunking strategies, embedding models, vector store configuration, and retrieval ranking - optimizing for clinical relevance, not just semantic similarity * Build a context assembly system that selects and structures the most relevant clinical information for each model invocation, given token constraints and clinical priority * Evaluate retrieval quality rigorously: build test sets, measure recall and precision, and iterate on the pipeline based on where retrieval fails Evaluation Framework Design * Design evaluation frameworks that assess AI recommendation quality beyond standard NLP metrics - working with clinical stakeholders to define what 'good' means for each use case * Build automated evaluation pipelines that can test AI outputs at scale: LLM-as-judge evaluators, human review workflows, and clinical validity checks * Maintain evaluation datasets that reflect the real distribution of clinical scenarios the model encounters in production * Report evaluation results in terms that clinical and product stakeholders can understand and act on Model Gap Analysis & Mitigation * Systematically identify where foundation model capabilities fall short for AnswersNow's care model: what clinical reasoning the model gets wrong, what it hallucinates, what it doesn't know how to handle * For each identified gap, recommend and implement the appropriate mitigation - improved retrieval, prompt engineering, output validation, or escalation to human review * Stay current on foundation model capabilities and evaluate new models against our clinical requirements as they emerge * Maintain a gap log and roadmap that gives product and clinical leadership visibility into current AI limitations and the plan to address them Production Monitoring & Quality * Monitor production AI outputs for quality, drift, and failure modes using the evaluation infrastructure you've built * Define alerting thresholds and escalation paths for when AI quality falls below acceptable clinical standards * Partner with the engineering team on observability - ensuring AI outputs are logged, traceable, and auditable * Conduct root-cause analysis when AI quality issues are reported and drive systematic fixes Clinical & Cross-Functional Partnership * Work closely with clinical leadership and BCBAs to understand the care model deeply enough to design AI systems that support it accurately * Translate clinical domain knowledge into technical requirements: what context does the model need, what outputs are clinically acceptable, where does the model need to defer to the clinician * Partner with the BI Engineer and data team on the data infrastructure that feeds the AI pipeline - session data, outcomes data, treatment plan content * Communicate AI system behavior clearly to non-technical stakeholders: what the system does, what it doesn't do, and where human judgment remains essential What we Offer * $144,000- $168,000 annual salary * Fully remote - work from anywhere in the U.S. * Flexible hours with an async-friendly team culture ## Related Videos - [Implementing continuous delivery in a data processing pipeline](https://www.wearedevelopers.com/videos/73-implementing-continuous-delivery-in-a-data-processing-pipeline) - [Introduction to Responsible AI: Balancing Value and Risk](https://www.wearedevelopers.com/videos/1972-introduction-to-responsible-ai-balancing-value-and-risk) - [How will artificial intelligence change the future of software testing?](https://www.wearedevelopers.com/videos/85-how-will-artificial-intelligence-change-the-future-of-software-testing) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [OLAP for AI Applications and why you should care](https://www.wearedevelopers.com/videos/100212-olap-for-ai-applications-and-why-you-should-care) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)