AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation

Apple Inc.
Cupertino, CA, United States
21 days ago
Apply on www.nexxt.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Image Quality Machine Learning Software Product Management Prompt Engineering Generative AI Siri Information Technology Machine Learning Operations Artificial Intelligence Markup Language (AIML)

Job description

Do you get excited by building AI applications to enhance the evaluation of various Apple AI products? Our Evaluation organization is responsible for providing principled assessments across a diverse range of Apple features, from Search and Siri to the latest Apple Intelligence capabilities. Within this critical function, our team specializes in leveraging advanced AI/ML techniques to enhance both the quality and efficiency of these comprehensive evaluations.

We are seeking a highly innovative and passionate Applied AI Scientist to develop cutting-edge AI/ML models for the automatic grading and quality assessment of our internal GenAI products., In this pivotal role, you will design and advance state-of-the-art autograder systems that evaluate various AI product quality at scale. You will apply deep expertise in prompt engineering, foundation model adaptation, and evaluation methodology to build robust, trustworthy, and extensible autograders that assess product performance, user experience, and adherence to quality and safety standards. Then you will collaborate with product, annotation, evaluation data scientists, autograder tooling engineers to deploy the state-of-the-art autograders, directly impacting the quality and success of Apple’s next-generation AI-powered features.

Requirements

  • Extensive experience with prompting techniques.
  • Deep understanding of GenAI models and 1+ year of industry experience in building or evaluation GenAI models.
  • Familiarity with LLMOps processes for deploying, monitoring and hillclimbing AI models in production environments.
  • Excellent analytical skills and judgement, capable of assessing data quality, diagnosing autograder limitations or biases, synthesizing findings into actionable insights, and communicating them clearly across teams.
  • MS/PhD degree in Computer Science, Machine Learning, AI, or a related fields.
  • Ownership mindset with the flexibility to take on whatever tasks-including annotation or operational work-are necessary to deliver results., * Experience in developing AI models specifically for quality assessment or automated feedback generation.
  • Familiarity with human annotation operations, sampling strategies, or subjective judgment evaluation.
  • Experience in designing human-in-the-loop evaluation workflows.
  • Familiar with image quality evaluation.
  • Demonstrated passion for leveraging AI to improve work efficiency and scale.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.nexxt.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:16 min

Evaluating the enduring financial and technological legacy of Apple

Marco Landi · World Congress 2024

2:37 min

Tracing the evolution from early AI to generative AI

Mike Mike · World Congress 2025

2:28 min

Why raw OCR pipelines fail in reality

Nazeer Saeed Nazeer Saeed · World Congress 2026 Europe

1:34 min

Introduction to the Apple Intelligence developer ecosystem

MIlan Todorović MIlan Todorović · World Congress 2025

2:14 min

Introduction to generative AI and content warnings

Cheuk Ho · World Congress 2023

40 sec

Navigating Apple's evolving on-device AI and machine learning stack

Precious Osaro Precious Osaro · World Congress 2026 Europe

Videos

See all

Related articles

See all