> Markdown version of [/jobs/ext/668108-applied-ai-scientist](https://www.wearedevelopers.com/jobs/ext/668108-applied-ai-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Applied AI Scientist - **Company:** Lever, Inc. - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $182,300.0 - $220,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Python (Programming Language), Machine Learning, Operational Databases, Performance Tuning, Standard Sql, Large Language Models - **Published:** June 27, 2026 - **Apply:** https://jobs.lever.co/ro/7d3b2a27-56e2-460a-9356-41022d9ed095 ## About the Role * 5+ years of experience in data science, applied machine learning, experimentation, or a closely related field, with at least the last year focused on applied LLMs or AI evaluation. * Strong Python and SQL skills with experience working on production data pipelines and experimentation. * You have experience designing reproducible evaluation frameworks rather than relying on manual spot checks or qualitative assessments. * You have strong statistical intuition: you think in terms of distributions, confidence intervals, variance, and sample sizes rather than anecdotes. * You're comfortable working closely with engineers and product teams to translate experimental findings into production improvements * Bonus: Experience with evaluation platforms (e.g. Braintrust, LangSmith, OpenAI Evals), experimentation platforms, causal inference, healthcare, or operations-heavy environments. ## Description Ro is building a team focused on shipping LLM-powered products across the patient experience, clinical operations, and internal tooling. We're hiring a Senior Applied AI Scientist to own the evaluation, measurement, and optimization of our AI systems. This role sits at the intersection of data science, applied machine learning, and product engineering. You'll design the frameworks that tell us whether our AI systems are actually working and use those insights to continuously improve them. This is not a research role. You'll work closely with engineers and product teams to evaluate production systems, run experiments, identify failure modes, and ensure our AI products become more accurate, reliable, and cost-effective over time., * Design and own evaluation frameworks for production LLM features, including LLM-as-a-judge evaluations, regression suites, synthetic datasets, golden datasets, and human review workflows. * Analyze production behavior to identify quality issues, hallucinations, latency bottlenecks, cost regressions, and emerging failure modes. * Design and run experiments including prompt variations, workflow changes, retrieval improvements, and model comparisons; and quantify their impact on quality, operational metrics, and user outcomes. * Define the metrics that matter and build dashboards that make AI performance visible across the organization. * Partner with engineering to determine which optimizations should be productionized and how to measure ongoing success. * Mentor teammates on experimental design, statistical rigor, evaluation methodology, and measurement best practices. ## Related Videos - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Practical performance tuning for Serverless Java on AWS](https://www.wearedevelopers.com/videos/2075-practical-performance-tuning-for-serverless-java-on-aws) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Analytics in the Age of Agentic AI: A tour of ClickHouse and Langfuse](https://www.wearedevelopers.com/videos/100240-analytics-in-the-age-of-agentic-ai-a-tour-of-clickhouse-and-langfuse) - [Three years of putting LLMs into Software - Lessons learned](https://www.wearedevelopers.com/videos/1508-three-years-of-putting-llms-into-software-lessons-learned) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)