Software Engineer Manager: Agentic Evaluation Platform

Apple Inc.
Cupertino, CA, United States
4 days ago
Apply on www.jobmonkeyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$237,600.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Big Data Cloud Computing Code Generation Continuous Integration Distributed Systems Python (Programming Language) Machine Learning Azure Machine Learning Management of Software Versions
+16 more
Data Logging Large Language Models Apache Spark Model Validation Siri Backend Kubernetes Information Technology Apache Flink Apache Kafka Data Management Grpc Data Pipelines Docker Golang Data Generation

Job description

Join the team redefining what a deeply personal and integrated assistant can be and lead the platform that defines how we measure it., In this role you’ll lead and grow a team of engineers building the backend services, compute orchestration, and evaluation harnesses that let us evaluate Siri reliably and at scale. You’ll own the technical direction and the roadmap, and you’ll be accountable for delivery.

The through-line of the work is interfaces. Your team owns the APIs that expose evaluation as a platform capability, the harness contracts that determine how devices and models are driven through a run, and the data workflows that carry results from execution through scoring to reporting. Those boundaries are the product: they decide how much compute the organization can actually use, how quickly a failure can be traced to its cause, and how many teams can build on evaluation without coming to you first. Get them right and the platform scales past your team; get them wrong and everything routes through you.

You’ll also have meaningful autonomy in how you get there. Both the evaluation platform and the underlying Siri architecture are evolving, so the specific frameworks, runtimes, and components will change over time. We’re looking for a leader who operates well in that ambiguity, someone who can reshape team scope deliberately as the domain matures, decide what to invest in versus retire, and keep the team’s momentum through the change.

Just as importantly, you’ll build the team itself: recruiting, mentoring, and developing engineers including senior ICs, with a focus on engineering rigor, technical depth, and clear ownership.

Responsibilities

Lead and grow a team of engineers building backend services, compute orchestration, and evaluation harnesses

Set the team’s technical direction and roadmap, and own delivery against it

Own the backend services and APIs that expose evaluation as a platform capability - job submission, execution, and results retrieval

Own compute orchestration across heterogeneous device and host fleets: scheduling, capacity, allocation, and the throughput and utilization targets that follow

Own harness engineering - the layer that drives devices and models through a run, captures output, and hands off to scoring

Define and steward the interfaces and contracts between harness, pipelines, and consuming teams, including versioning, backward compatibility, and migration paths

Own the data workflows that move results from execution through scoring to storage and reporting, and the schemas they depend on

Deliver and evolve reliability and health dashboards and release-readiness signals that leadership relies on for ship decisions

Set the bar for diagnosability across the stack - from environment provisioning through pipeline execution to scoring - prioritizing automated triage and durable fixes over recurring manual investigation

Partner with QE engineers, feature teams, architecture leads, and infrastructure teams to align on interfaces, priorities, and shared standards, and to hold a defensible quality bar

Recruit, mentor, and develop engineers, including senior ICs, with a focus on engineering rigor, technical depth, and clear ownership

Operate effectively in ambiguity, reshaping team scope as the evaluation platform and the underlying Siri architecture evolve

Requirements

5+ years of experience developing production software (e.g., Python, Java, Go, or Swift)

3+ years of engineering management experience, including owning roadmap and delivery, mentoring, setting clear expectations, and providing consistent, honest feedback

Strong technical leadership and the ability to set architectural direction and inspire a team

Experience with test infrastructure, CI/CD, or evaluation/ML platforms

Proactive and self-motivated, with demonstrated creative and critical-thinking abilities

Excellent spoken and written communication skills, and the ability to collaborate across teams

M.S. or B.S. in Computer Science, Machine Learning, or a related field, or equivalent experience

Preferred Qualifications

Experience building or leading teams that apply LLMs / AI agents to real engineering workflows (e.g., data generation, code generation, or agentic pipelines with human-in-the-loop review)

Depth in one or more of: distributed systems and backend services (REST/gRPC), cloud infrastructure (Docker/Kubernetes), data platforms, or large-scale test and release infrastructure

Familiarity with LLM and Agentic evaluation concepts

Experience implementing model evaluation frameworks and experiment workflows, including offline and online experiments, to measure how model or platform changes affect key metrics such as accuracy, latency, and engagement

Experience building machine learning platforms and infrastructure used by multiple teams for experimentation and deployment

Experience designing and integrating monitoring, logging, and alerting solutions - dashboards, anomaly detection, and real-time alerts - to track reliability and performance in production systems

Experience developing data processing pipelines with big data frameworks (such as Apache Spark, Kafka, or Flink) to ingest, transform, and analyze large-scale log or event data

Demonstrated cross-functional leadership and the ability to drive coverage and quality decisions with senior stakeholders

Experience hiring and growing a team, including senior engineers

Comfort operating in ambiguity and reshaping team scope as the domain matures

Benefits & conditions

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $237,600 and $401,700, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

About the company

As part of the Siri Agentic Evaluation Engineering organization, you will help shape one of the world’s most widely used AI assistants, powered by our next generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS.

As the engineering leader of the evaluation platform and services, you’ll own the systems that run and scales Siri’s evaluations, leveraging the best Apple platform capabilities and unlocking new ones to create novel evaluations. Your team applies large scale infrastructure, apple platforms and service, and AI agents to steer, validate, and judge evaluation scenarios at scale.

These generate the signals that leadership relies on to make ship decisions. When we say a capability is ready, it’s your platform that backs the claim.

This is a rare opportunity to build scalable software and services for real world products at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:16 min

Evaluating the enduring financial and technological legacy of Apple

Marco Landi · World Congress 2024

4:56 min

Establishing internal service communication with gRPC

Florian Bader Florian Bader · World Congress 2026 Europe

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

2:23 min

Historical breakthroughs in natural language processing models

Mary Grygleski Mary Grygleski · LIVE

Videos

See all

Related articles

See all