> Markdown version of [/jobs/ext/2065634-data-scientist-agent](https://www.wearedevelopers.com/jobs/ext/2065634-data-scientist-agent). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist, Agent - **Company:** Lovable Labs Incorporated - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Data Analysis, BigQuery, Python (Programming Language), Standard Sql, SQL Databases, Google Cloud, Large Language Models, Virtual Agents - **Published:** August 15, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/pb8wv8acj7 ## About the Role * A data scientist who wants to make an AI agent measurably better, not just report on it. You own agent quality metrics and drive them up. * Experience or strong interest in LLM evaluation and observability: building evals, scoring outputs, tracing agent behavior, and catching regressions. * Strong SQL and Python, applied statistics, and experimentation. Comfortable designing A/B tests for agent changes where outcomes are noisy. * You build systems and agents that produce this insight continuously, rather than one-off analyses. * Instinct for what "good" looks like in agent behavior (success, error rates, task completion) and how to measure it when there is no clean answer key. * Entrepreneurial. Thrives in ambiguity, and works closely with the engineers building the agent., Please submit your application in English. It's our company language, so you'll be speaking lots of it if you join. ## Description * Define and own the metrics for agent quality: success, completion, error rates, and the behaviors that drive them. * Build the eval systems and experiment framework that decide whether an agent change ships, like an A/B-tested rollout that catches a change increasing errors before it reaches everyone. * Turn agent traces and telemetry into concrete fixes, working directly with the agent engineering team. * Build the tooling and agents that produce these evaluations continuously as the agent evolves. * Set the bar for how we judge agent behavior where there is no answer key to check against. Our tech stack We're building with tools that both humans and AI love: * Languages: SQL and Python * LLM evaluation & observability: Braintrust, OTEL tracing, many LLM providers * Warehouse & events: BigQuery, PubSub * Analytics & product: Hex, Lovable Apps * Experimentation: A/B and growth testing * Cloud: GCP And always on the lookout for what's next. How We Hire * Fill in a short form and jump on an intro call with our recruiting team * A call with the hiring manager * A take-home case study * A Most Impressive Project session * Cross-functional interviews with the people you'd work with * A final conversation with leadership ## Related Videos - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) - [AI Agents & Agentic AI](https://www.wearedevelopers.com/videos/2017-ai-agents-agentic-ai) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Analytics in the Age of Agentic AI: A tour of ClickHouse and Langfuse](https://www.wearedevelopers.com/videos/100240-analytics-in-the-age-of-agentic-ai-a-tour-of-clickhouse-and-langfuse) - [Making Data Warehouses fast. A developer's story.](https://www.wearedevelopers.com/videos/302-making-data-warehouses-fast-a-developer-s-story) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering) - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [The Overflow: AI and Agentic Coding](https://www.wearedevelopers.com/magazine/721-the-overflow-ai-and-agentic-coding)