AIML - Applied AI Scientist, Image Autograder Systems, Evaluation

Apple Inc.
Cupertino, CA, United States
26 days ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Image Quality Python (Programming Language) Machine Learning Language Modeling Generative AI Information Technology Artificial Intelligence Markup Language (AIML)

Job description

We are looking for an AI Scientist to join a centralized evaluation organization building the next generation of autograders across Apple’s most visible image generation AI features. In this role, you will develop autograders that reliably score image output quality at scale. Your work directly influences the quality of AI experiences used by hundreds of millions of Apple customers.

This is a high-impact individual contributor role at the intersection of genAI evaluation, genAI model training, data quality, and AI engineering. You will work closely with senior AI scientists, MLEs, data annotation teams, and feature engineers in a fast-moving, technically rigorous environment., In this role you will focus on Autograder research, training and adoption. Collaborate with AI feature teams, eval design teams, and annotation teams to refine feature requirements, grading rubrics and gold annotation sets. Develop, evaluate, and iterate on grading prompts to align autograder behavior with grading rubrics and the gold sets. Identify when prompt tuning reaches its limits and apply other advanced techniques such as fine-tuning to close remaining grading accuracy gaps. Design insightful analysis to measure and explain Autograder quality. Collaborate with MLEs and feature teams on autograder deployment and adoption. Develop scalable system to speed up autograder training processes.

Requirements

  • Master’s or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
  • Deep understanding of visual-language models.
  • Familiarity with image quality assessment - perceptual quality dimensions, evaluation metrics, and rubric design.
  • Proficiency in Python; capable of writing well-structured, production-ready model code.
  • Strong communication skills in explaining autograder quality and driving autograder adoption with partner teams., * 1+ years of industry experience in building VLM-based products.
  • Familiarity with autograder and evaluator-specific concepts: grading accuracy, agreement with human raters, calibration, and rubric design.
  • Strong expertise in prompt tuning and fine tuning for VLMs.
  • Demonstrated ability to read AI literature and translate it into applied autograder experiments.
  • Prior experience in building agentic system to scale autograder training/validation.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Introduction to the Apple Intelligence developer ecosystem

MIlan Todorović MIlan Todorović · World Congress 2025

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

2:37 min

Tracing the evolution from early AI to generative AI

Mike Mike · World Congress 2025

2:28 min

Why raw OCR pipelines fail in reality

Nazeer Saeed Nazeer Saeed · World Congress 2026 Europe

40 sec

Navigating Apple's evolving on-device AI and machine learning stack

Precious Osaro Precious Osaro · World Congress 2026 Europe

1:57 min

Evolution of machine learning algorithms and computing hardware

Alexandra Waldherr · LIVE

Videos

See all

Related articles

See all