Data Scientist, Agent

Lovable Labs Incorporated
United States
22 days ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Job source

Tech stack

A/B Testing Artificial Intelligence Data Analysis BigQuery Python (Programming Language) Standard Sql SQL Databases Google Cloud Large Language Models Virtual Agents

Job description

  • Define and own the metrics for agent quality: success, completion, error rates, and the behaviors that drive them.
  • Build the eval systems and experiment framework that decide whether an agent change ships, like an A/B-tested rollout that catches a change increasing errors before it reaches everyone.
  • Turn agent traces and telemetry into concrete fixes, working directly with the agent engineering team.
  • Build the tooling and agents that produce these evaluations continuously as the agent evolves.
  • Set the bar for how we judge agent behavior where there is no answer key to check against.

Our tech stack

We’re building with tools that both humans and AI love:

  • Languages: SQL and Python
  • LLM evaluation & observability: Braintrust, OTEL tracing, many LLM providers
  • Warehouse & events: BigQuery, PubSub
  • Analytics & product: Hex, Lovable Apps
  • Experimentation: A/B and growth testing
  • Cloud: GCP

And always on the lookout for what’s next.

How We Hire

  • Fill in a short form and jump on an intro call with our recruiting team
  • A call with the hiring manager
  • A take-home case study
  • A Most Impressive Project session
  • Cross-functional interviews with the people you’d work with
  • A final conversation with leadership

Requirements

  • A data scientist who wants to make an AI agent measurably better, not just report on it. You own agent quality metrics and drive them up.
  • Experience or strong interest in LLM evaluation and observability: building evals, scoring outputs, tracing agent behavior, and catching regressions.
  • Strong SQL and Python, applied statistics, and experimentation. Comfortable designing A/B tests for agent changes where outcomes are noisy.
  • You build systems and agents that produce this insight continuously, rather than one-off analyses.
  • Instinct for what “good” looks like in agent behavior (success, error rates, task completion) and how to measure it when there is no clean answer key.
  • Entrepreneurial. Thrives in ambiguity, and works closely with the engineers building the agent., Please submit your application in English. It’s our company language, so you’ll be speaking lots of it if you join.

About the company

At Lovable, data scientists are not isolated model-builders; they sit close to the product, experimenting continuously with how intelligence changes user behavior and product dynamics.

Why Lovable?

Lovable lets anyone and everyone build software with any language. From solopreneurs to Fortune 100 teams, millions of people use Lovable to transform raw ideas into real products, fast. We are at the forefront of a foundational shift in software creation, which means you have an unprecedented opportunity to change the way the digital world works. Lovable-built applications and websites are visited hundreds of millions of times a month, and our enterprise footprint is compounding fast. And we’re just getting started.

We’re a small, talent-dense team building a generation-defining company from Stockholm. We value extreme ownership, high velocity, and low-ego collaboration. We seek out people who care deeply, ship fast, and are eager to make a dent in the world.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:20 min

Fostering curiosity through accessible no-code AI tools

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

4:06 min

Using ClickHouse as a foundation for fast analytics

Hellmar Becker Hellmar Becker · World Congress 2026 Europe

54 sec

Generating multiple hook options for outreach A/B testing

Leandro Gomes da Silva Leandro Gomes da Silva · World Congress 2025

2:30 min

Leveraging BigQuery ML for scalable SQL-based segmentation experiments

Julian Joseph · LIVE

56 sec

Performing local A/B testing across multiple AI agents

Julia Kasper · Coffee With Developers

1:56 min

Creating immutable blockchain tables using standard SQL syntax

Wei Hu Wei Hu · World Congress 2023

Videos

See all

Related articles

See all