AIML - Software Engineer - AI, Evaluation
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+5 more
Job description
Do you get excited by building software systems to enhance the automatic evaluation of various Apple AI products? Our Evaluation organization is responsible for providing principled assessments across a diverse range of Apple features, from Search, Siri to the latest Apple Intelligence capabilities. Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer to expand our tools and systems., As an AI Software Engineer on the team, you will design and build tools and systems that sit at the intersection of AI modeling, software engineering, and product quality. You will design and develop extensible frameworks, pipelines, and tools that enable efficient development, deployment, and qualitative measurement of AI models. Due to the breadth of products supported, the role requires strong software design and engineering skills. Your work will directly influence product launch decisions and enable teams across Apple to iterate faster and with greater confidence.
Requirements
- BS/MS/PhD degree in Computer Science, Machine Learning, AI, or a related field.
- Exceptional Python skills.
- Solid software engineering fundamentals with production experience, including system design, API design, CI/CD, testing strategies, code maintainability, system monitoring, debugging complex systems and etc.
- Demonstrated expertise in using AI-assisted software development workflows to accelerate software development while maintaining code quality.
- Strong communication skills and proven ability to work collaboratively with cross-functional teams., * Experience with building LLM applications, frameworks, and offline evaluations.
- Familiar with MLOps principles for model lifecycle management.
- Experience in building scalable tools for product quality evaluation.
- Ability to understand and interpret evaluation reports, including metrics such as precision, recall, run-to-run consistency, and common pitfalls like data leakage.
- Product-minded, with a strong ability to translate ambiguous product requirements into solutions.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Navigating the AI Shift
What Are Large Language Models?
Dev Digest 137 - AI'm not sure about this