Staff Experience Designer, AI Evaluation Platform

Apple Inc.
New York, NY, United States
8 days ago
Apply on www.jobmonkeyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$175,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Design of User Interfaces Machine Learning SQL Databases

Job description

Teams building AI features need to know if what they shipped works. Today that means navigating unfamiliar territory: scoring something non-deterministic, telling a real regression from noise, trusting a judgment a model made instead of a person. Most existing tools here were built for the researchers who invented these methods, not for the people who now need to use them daily.

You’ll set the direction here and build the design system everyone else builds on. We’d rather get a rough version in front of real users than a polished one later. And because teams use this platform to decide what to ship, trust matters more than usual: provenance, uncertainty, and edge cases determine whether someone acts on a result or quietly stops believing it.

What makes this team unusual is its interdisciplinary core. Alongside the platform, we run a research group working on evaluation methodology itself, including how to tell whether an evaluator is calibrated, biased, or measuring what it claims to. Their methods ship into this platform, and you decide how anyone first encounters them. You’ll be close to that work while it’s still forming, instead of picking it up once it’s finished.

Responsibilities

Design the path from a vague question about a system’s quality or safety to a running evaluation, without requiring evaluation expertise.

Make dense, multi-dimensional results legible enough to act on: scores, comparisons, distributions, and the uncertainty attached to them.

Define how someone inspects multi-step agent behavior, traces a failure to its cause, and audits a judgment a model made.

Partner with our research scientists as new evaluation techniques are being developed, so you shape how they surface in the product instead of designing around them after they’re set. You’ll translate these techniques into interfaces non-specialists can use and trust.

Build and maintain the platform’s design system as the first designer on the team.

Prototype against real data to validate ideas before they’re built.

Run research directly with the engineers, scientists, and practitioners who use the platform.

Requirements

8+ years of experience designing digital products or experiences, including end-to-end ownership of complex, data-dense products for technical users. You’ve set direction in spaces with no existing precedent, and you get to clarity by running research yourself with the people who use what you build.

A portfolio you can share, including at least one case study that walks through your process from problem framing to shipped outcome.

Proven ability to design complex information: dashboards, comparison views, large result sets, and results that carry statistical uncertainty.

Strong interaction and visual design skills, with a high bar for craft in dense, information-heavy interfaces.

Enough fluency in AI/ML concepts (benchmarks, metrics, model-based judging, agentic systems) to work directly with a research scientist or engineer.

Fluency with AI tools in your own practice. You use Claude Code or equivalents to build working prototypes and extend what you can make on your own.

Excellent communication skills, with the ability to build buy-in across engineering and research without formal authority.

Preferred Qualifications

Experience as the first or only designer on a platform, and/or experience translating research output into shipped product.

Experience building and maintaining a design system, including its components, patterns, and naming conventions.

An AI-first instinct for interface design: shipped interfaces organized around stated intent, or surfaces meant to be operated by agents as well as people.

Experience designing coherent workflows across multiple surfaces (UI, CLI, SDK) that need to stay consistent with each other.

Familiarity with modern evaluation, observability, or visualization tooling (e.g. LangSmith, Braintrust, D3, Vega-Lite).

Comfort with SQL and notebooks to explore your own data.

Benefits & conditions

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $175,000 and $308,500, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

About the company

AI systems are only as trustworthy as the methods used to evaluate them. At Apple, where AI powers experiences for billions of people, getting evaluation right is not a support function, it is a foundational science.

Our team, part of Apple Services Engineering, builds the platform that teams across Apple use to evaluate the AI and agentic systems they ship. It’s where they define what "good" means, prove it, and act on what they find. Evaluating non-deterministic systems is one of the hardest unsolved problems in production ML, and one Apple has to get right at scale.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:16 min

Adjusting technical interviews for an AI-native industry

Ayotunde Obasa Ayotunde Obasa · Europe 2026 Virtual

1:14 min

Evolution of distributed SQL database architectures

Wei Hu Wei Hu · World Congress 2024

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

3:46 min

Core terminology and audiences for interpretable artificial intelligence

Karol Przystalski · LIVE

5:33 min

Applying product thinking to internal platform design

Christian Strack · LIVE

1:36 min

Evolution from key-value stores to distributed SQL

Wei Hu Wei Hu · World Congress 2025

Videos

See all

Related articles

See all