> Markdown version of [/jobs/ext/2844266-sr-machine-learning-engineer-speech-llm-evaluation](https://www.wearedevelopers.com/jobs/ext/2844266-sr-machine-learning-engineer-speech-llm-evaluation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Machine Learning Engineer, Speech LLM Evaluation - **Company:** Apple Inc. - **Location:** Cupertino, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Distributed Computing Environment, Python (Programming Language), Machine Learning, Large Language Models, Apache Spark, Model Validation, Information Technology, Data Pipelines - **Published:** September 11, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3386409878&tx=CT3432TYD&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience. * Experience building or working with text, speech or audio evaluation pipelines and metrics. * Proficiency in Python and experience building data processing pipelines at scale. * Experience curating or annotating datasets for machine learning evaluation or training. * Working knowledge of statistics as applied to measuring model performance and interpreting evaluation results. * Familiarity with large language model evaluation techniques, including automated (LLM-as-judge) and human evaluation methods. * Strong written and verbal communication skills, with the ability to explain evaluation results to both technical and non-technical audiences. Preferred Qualifications * Experience evaluating audio-native or multimodal (speech-in, speech-out) large language models. * Experience designing or running human evaluation studies (e.g., side-by-side comparisons, MOS ratings) at scale. * Familiarity with personalization and named-entity evaluation challenges in speech systems. * Experience with multilingual or international audio dataset development. * Experience with distributed data processing frameworks (e.g., Spark) for large-scale audio dataset generation. * Publication record or demonstrated contributions in speech, audio ML, or NLP evaluation. ## Description This role owns the data and metrics foundation for evaluating speech LLMs (e.g., real-time speech understanding and generation models) across accuracy, robustness, and conversational quality. You'll build and curate evaluation datasets that reflect real usage - from personalized named-entity queries to multi-turn fluid conversations - and design the metrics and automated judges that turn model outputs into actionable, trustworthy signal. You'll work closely with modeling, infrastructure, and product partners to make sure every new model is evaluated quickly, consistently, and at the right level of rigor before it reaches customers. ## Related Videos - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Inside the Mind of an LLM](https://www.wearedevelopers.com/videos/1617-inside-the-mind-of-an-llm) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Implementing continuous delivery in a data processing pipeline](https://www.wearedevelopers.com/videos/73-implementing-continuous-delivery-in-a-data-processing-pipeline) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)