Data Scientist

Kaizo
Amsterdam, Netherlands
29 days ago

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
0 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Data Analysis Audio Signal Processing BigQuery Software as a Service Elasticsearch Python (Programming Language) MongoDB NumPy Operational Databases SQL Databases Google Cloud
+10 more
Large Language Models Jupyter Pandas Kubernetes Apache Kafka Machine Learning Operations Front End Software Development Stream Processing Docker Microservices

Job description

Do you want to work out how you actually measure whether an AI system is doing a good job, and then make it better? We’re looking for a sharp, analytical data scientist to own the evaluation and quality loop of our AutoQA product. You can be early in your career (recent graduates with strong LLM fundamentals are welcome), we have room and a clear path for you to grow.

In a nutshell o Join a fast-growing SaaS company in an international environment (steep learning curve guaranteed). o Own a high-impact problem: making LLM-powered quality assurance measurably accurate at scale. o Sit at the intersection of AI and CX, working directly with enterprise customers and their real-world QA rubrics. o Grow into a Senior Data Scientist, AI Engineer, or ML Engineer role. We invest in progression. o Enjoy the perks: flexible hours, open holiday policy, an office in the heart of Amsterdam with hybrid flexibility, visa sponsorship, great gear, workations, and team events.

About Kaizo At Kaizo, we build a performance development and quality platform for customer support teams. Our AutoQA product uses LLMs to review support conversations against each customer’s own quality rubric, automatically and at scale. Behind it sits a microservices-based stream processing platform handling over 200 million events per day (Kafka, Kubernetes on Google Cloud, ElasticSearch, MongoDB, BigQuery), and an LLMOps stack built around LangSmith for experimentation, prompt management, and tracing. The hard part isn’t calling an LLM. It’s knowing, with evidence, how well the system performs on every customer’s unique rubric, and having a reliable, repeatable way to improve it. That’s where you come in.

What you’ll focus on o Translate customer rubrics into AutoQA instructions. Work with real customer quality criteria and turn them into precise, testable instructions that LLMs can score reliably. o Run experiments that move accuracy. Design and execute evaluation experiments on large, representative datasets using LangSmith and BigQuery, and track quality with our performance metrics. o Build the datasets that make evaluation possible. Curate raw production data into golden datasets with balanced coverage, and generate synthetic data to cover the rare cases that matter most. QA is a discipline of rare occurrences: distributions are skewed, positives are scarce, and resourcefulness beats volume. o Build LLM-as-a-judge pipelines to assess system quality internally and make evaluation repeatable. o Make results actionable. Your experiments should end in a decision: change a prompt, adjust which tools the system uses, surface context the AI is missing, or flag where new capabilities are needed. You’ll work with the AI team to ship those decisions. o Evaluate across the full stack. Beyond scoring quality, you’ll help validate retrieval (RAG/IR), tool calling, and speech pipelines (transcription and diarization quality). o Join customer calls with the team to understand how QA leaders define quality, and feed what you learn back into the product.

What you’ll grow into o Shaping how customers monitor quality themselves, catch drift, and keep their AutoQA setup improving over time. o Smarter categorization and routing of conversations to power analytics and get the right tickets to the right evaluation., You’ll join our AI team (two data scientists, an AI engineer, and an ML engineer) and collaborate closely with our data engineers, frontend engineers, designer, product manager, and CX teams. You’ll have mentorship from day one and real ownership fast.

Requirements

o A solid understanding of how LLMs work and hands-on experience prompting them for accuracy (coursework, thesis, internships, or side projects all count; production experience is a bonus). o Good applied statistics: experiment design, classifier evaluation, precision/recall trade-offs, and working with heavily imbalanced data. o Strong Python skills and fluency with the standard data toolkit (Pandas, NumPy, Jupyter). SQL is a plus. o An analytical, evidence-first mindset: you’d rather measure than assume. o Product sense and empathy for end users. You’ll be building for QA managers and support agents, not just for benchmarks. o Excellent written and verbal communication. You’ll present findings to the team and join customer conversations. o 0 to 2 years of industry experience. Recent graduates with strong relevant work are encouraged to apply. o A team player who’s comfortable wearing multiple hats. We’re an early-stage company and things move fast. Bonus points for: o Experience with LangSmith or similar LLMOps/evaluation tooling o Google Cloud Platform, BigQuery, Docker, or Kubernetes o Speech/audio processing or ASR evaluation o Fine-tuning or building synthetic datasets for LLMs

Benefits & conditions

What’s in it for you? o An office in the heart of Amsterdam, with the flexibility of hybrid working o Visa sponsorship available for eligible candidates o Great office gear: MacBook, tools, desk, chair, whatever you need o Flexible working schedule and an open holiday policy o Fun workations and team events o A clear growth path into Senior Data Scientist, AI Engineer, or ML Engineer roles or

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on kaizo.recruitee.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

1:35 min

How LLMs imitate human data discovery behavior

Jordan Tigani Jordan Tigani · World Congress 2026 Europe

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · World Congress 2025

3:23 min

Exploring specialized career paths within the data science ecosystem

Julian Joseph · LIVE

Videos

See all

Related articles

See all