> Markdown version of [/jobs/ext/3029275-software-engineer-manager-agentic-evaluation-platform](https://www.wearedevelopers.com/jobs/ext/3029275-software-engineer-manager-agentic-evaluation-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer Manager: Agentic Evaluation Platform - **Company:** Apple Inc. - **Location:** Cupertino, CA, United States - **Experience:** Experienced - **Salary:** $237,600.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Big Data, Cloud Computing, Code Generation, Continuous Integration, Distributed Systems, Python (Programming Language), Machine Learning, Azure Machine Learning, Management of Software Versions, Data Logging, Large Language Models, Apache Spark, Model Validation, Siri, Backend, Kubernetes, Information Technology, Apache Flink, Apache Kafka, Data Management, Grpc, Data Pipelines, Docker, Golang, Data Generation - **Published:** September 22, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/28038671/Software-Engineer-Manager-Agentic-Evaluation-Platform-California-Cupertino-7413 ## About the Role 5+ years of experience developing production software (e.g., Python, Java, Go, or Swift) 3+ years of engineering management experience, including owning roadmap and delivery, mentoring, setting clear expectations, and providing consistent, honest feedback Strong technical leadership and the ability to set architectural direction and inspire a team Experience with test infrastructure, CI/CD, or evaluation/ML platforms Proactive and self-motivated, with demonstrated creative and critical-thinking abilities Excellent spoken and written communication skills, and the ability to collaborate across teams M.S. or B.S. in Computer Science, Machine Learning, or a related field, or equivalent experience Preferred Qualifications Experience building or leading teams that apply LLMs / AI agents to real engineering workflows (e.g., data generation, code generation, or agentic pipelines with human-in-the-loop review) Depth in one or more of: distributed systems and backend services (REST/gRPC), cloud infrastructure (Docker/Kubernetes), data platforms, or large-scale test and release infrastructure Familiarity with LLM and Agentic evaluation concepts Experience implementing model evaluation frameworks and experiment workflows, including offline and online experiments, to measure how model or platform changes affect key metrics such as accuracy, latency, and engagement Experience building machine learning platforms and infrastructure used by multiple teams for experimentation and deployment Experience designing and integrating monitoring, logging, and alerting solutions - dashboards, anomaly detection, and real-time alerts - to track reliability and performance in production systems Experience developing data processing pipelines with big data frameworks (such as Apache Spark, Kafka, or Flink) to ingest, transform, and analyze large-scale log or event data Demonstrated cross-functional leadership and the ability to drive coverage and quality decisions with senior stakeholders Experience hiring and growing a team, including senior engineers Comfort operating in ambiguity and reshaping team scope as the domain matures ## Description Join the team redefining what a deeply personal and integrated assistant can be and lead the platform that defines how we measure it., In this role you'll lead and grow a team of engineers building the backend services, compute orchestration, and evaluation harnesses that let us evaluate Siri reliably and at scale. You'll own the technical direction and the roadmap, and you'll be accountable for delivery. The through-line of the work is interfaces. Your team owns the APIs that expose evaluation as a platform capability, the harness contracts that determine how devices and models are driven through a run, and the data workflows that carry results from execution through scoring to reporting. Those boundaries are the product: they decide how much compute the organization can actually use, how quickly a failure can be traced to its cause, and how many teams can build on evaluation without coming to you first. Get them right and the platform scales past your team; get them wrong and everything routes through you. You'll also have meaningful autonomy in how you get there. Both the evaluation platform and the underlying Siri architecture are evolving, so the specific frameworks, runtimes, and components will change over time. We're looking for a leader who operates well in that ambiguity, someone who can reshape team scope deliberately as the domain matures, decide what to invest in versus retire, and keep the team's momentum through the change. Just as importantly, you'll build the team itself: recruiting, mentoring, and developing engineers including senior ICs, with a focus on engineering rigor, technical depth, and clear ownership. Responsibilities Lead and grow a team of engineers building backend services, compute orchestration, and evaluation harnesses Set the team's technical direction and roadmap, and own delivery against it Own the backend services and APIs that expose evaluation as a platform capability - job submission, execution, and results retrieval Own compute orchestration across heterogeneous device and host fleets: scheduling, capacity, allocation, and the throughput and utilization targets that follow Own harness engineering - the layer that drives devices and models through a run, captures output, and hands off to scoring Define and steward the interfaces and contracts between harness, pipelines, and consuming teams, including versioning, backward compatibility, and migration paths Own the data workflows that move results from execution through scoring to storage and reporting, and the schemas they depend on Deliver and evolve reliability and health dashboards and release-readiness signals that leadership relies on for ship decisions Set the bar for diagnosability across the stack - from environment provisioning through pipeline execution to scoring - prioritizing automated triage and durable fixes over recurring manual investigation Partner with QE engineers, feature teams, architecture leads, and infrastructure teams to align on interfaces, priorities, and shared standards, and to hold a defensible quality bar Recruit, mentor, and develop engineers, including senior ICs, with a focus on engineering rigor, technical depth, and clear ownership Operate effectively in ambiguity, reshaping team scope as the evaluation platform and the underlying Siri architecture evolve ## Related Videos - [Building a Multi-Agent Orchestration Engine That Actually Follows the Rules](https://www.wearedevelopers.com/videos/100159-building-a-multi-agent-orchestration-engine-that-actually-follows-the-rules) - [Exploring the Power of gRPC-Gateway for Writing RESTful Services](https://www.wearedevelopers.com/videos/2072-exploring-the-power-of-grpc-gateway-for-writing-restful-services) - [Lessons from Steve Jobs - Learnings from the Past for the Future](https://www.wearedevelopers.com/videos/1021-lessons-from-steve-jobs-learnings-from-the-past-for-the-future) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Evals vs. Evil - AI and Package Security - Laurie Voss](https://www.wearedevelopers.com/videos/2131-evals-vs-evil-ai-and-package-security-laurie-voss) - [Boosting OpenSearch Performance: gRPC Search in Action](https://www.wearedevelopers.com/videos/1964-boosting-opensearch-performance-grpc-search-in-action) ## Related Articles - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)