Annotation Data Scientist

Apple Inc.
Washington, United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Cognitive Science Data Visualization Human-Computer Interaction Machine Learning Natural Language Processing SQL Databases Large Language Models Multi-Agent Systems Prompt Engineering Apache Spark Jupyter Siri
+2 more
Pandas Information Technology

Job description

Play a part in the ongoing revolution in human-computer interaction. Siri is evolving - and the way we evaluate it has to evolve with it. Join the Evaluation Integrity team to help build the trusted quality signal behind every Siri release. Within the Siri evaluation organization, the Human Evaluation sub-team is responsible for answering the question: can we trust our evals? We do that by designing human-in-the-loop (HITL) annotation tasks that scrutinize every moving part of an agentic evaluation - the simulated user agent, the conversation it has with Siri, and the automated evaluators that grade the exchange. This role sits at the intersection of data science, human annotation engineering, and evaluation methodology, and is instrumental in turning human judgment into a rigorous, reproducible signal that directly informs pre-ship model and product decisions. As an Annotation Data Scientist on the Evaluation Integrity team, you will design and run HITL annotation projects that evaluate the quality and authenticity of agentic user personae, the validity of agent-to-agent conversations, and the reliability of LLM-as-judge and rule-based evaluators against Siri’s product specifications. You will own annotation initiatives end-to-end; from rubric design and tooling, through annotator calibration, to data science analysis that turns annotator judgments into actionable signal for modeling, planning, and product teams.

Requirements

Bachelor’s or Master’s degree in a quantitative or related field such as Data Science, Computer Science, Linguistics, Statistics, or Cognitive Science, or equivalent job-related experience.

5+ years of hands-on experience working with human-annotated datasets or human-in-the-loop evaluation methodologies for machine learning, natural language processing, or large language model systems.

5+ years of experience using Python for data processing, analysis, and prototyping, including experience with libraries such as pandas, Jupyter, and at least one data visualization library.

Experience designing, implementing, and communicating annotation schemas, rubrics, or ontologies for machine learning training or evaluation data.

Experience managing multiple concurrent dataset curation efforts, including scoping work, iterating on guidelines, coordinating with in-house or vendor annotators, and monitoring annotator performance metrics such as accuracy, throughput, and inter-annotator agreement.

Experience specifying or designing custom annotation tooling in collaboration with software engineers.

Preferred Qualifications

Experience evaluating LLM-powered or agentic systems, including familiarity with LLM-as-judge methodologies, rubric-based grading, or trajectory and tool-call evaluation.

Familiarity with statistical methods that address accuracy and variability in human annotation data, such as inter-annotator agreement, Cohen’s or Fleiss’ kappa, Krippendorff’s alpha, or bootstrapping.

Data-querying experience with SQL, Spark, or similar, and comfort working with large, complex, real-world datasets.

Experience building pre-ship evaluation pipelines for conversational or assistant products.

Experience with prompt engineering, or with designing simulated user personae for agent evaluation.

Experience running annotation programs across multiple locales or at large scale.

Excellent written and verbal communication skills, with the ability to explain technical topics clearly to data scientists, engineers, annotators, and cross-functional partners.

Proven ability to collaborate effectively across functions and drive projects of varying sizes and scopes - knowing when to dive deep and when to delegate.

About the company

Join the team redefining what a deeply personal and integrated assistant can be. As part of the Siri organization, you will help shape one of the world’s most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS. This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:31 min

Essential AI and human skills for future teams

Alexander Weißhaupt Alexander Weißhaupt +1 · World Congress 2025

3:31 min

Producing candlestick visualization charts inside integrated Jupyter notebooks

Akmal Chaudhri Akmal Chaudhri · LIVE

1:16 min

Evaluating the enduring financial and technological legacy of Apple

Marco Landi · World Congress 2024

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

5:36 min

Building a data analysis stack with Python and Jupyter

Markus Harrer Markus Harrer · World Congress 2021

9:53 min

Developing programmatic training workflows using Python Jupyter Notebooks

Jose Luis Latorre Millas · LIVE

Videos

See all

Related articles

See all