> Markdown version of [/jobs/ext/1212947-ml-engineer-llm-evaluation-automation](https://www.wearedevelopers.com/jobs/ext/1212947-ml-engineer-llm-evaluation-automation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Engineer - LLM Evaluation & Automation - **Company:** Ririo.Com, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Big Data, Computer Programming, Data Normalization, Python (Programming Language), Machine Learning, Pattern Recognition, SQL Databases, Large Language Models, Prompt Engineering, Build Management, Pyspark, Information Technology - **Published:** July 9, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=470e82532c33762a ## About the Role We are seeking a highly skilled Machine Learning Engineer who specializes in leveraging Large Language Models (LLMs) for automated evaluation and quality assessment. In this role, you will design and build systems that automatically measure and improve the accuracy, relevance, and consistency of model outputs. You will lead initiatives to create evaluation pipelines, develop metrics, and deliver actionable insights for continuous improvements. This position requires strong technical expertise, analytical problem-solving abilities, and the capacity to manage projects across multiple cross-functional teams., 5+ years of experience in ML engineering, NLP, or AI/ML automation. * Hands-on experience in prompt engineering and designing LLM-based evaluation systems is preferred * Strong understanding of machine learning principles with focus on NLP and advanced LLM capabilities (e.g., Chain-of-Thought, agentic workflows) * Expertise in building automated evaluation or QA pipelines. * Excellent analytical and problem-solving skills with experience in root cause and error pattern analysis. * Proven project management and cross-functional collaboration experience. * Excellent communication skills to convey complex insights to technical and non-technical audiences. * Detail-oriented mindset with a focus on evaluation metrics, prompt design, and automation. * Ability to quickly adapt to new business rules and evaluation guidelines across diverse product domains. * Strong programming skills in Python and SQL. * Experience with big data technologies like PySpark for data aggregation and sampling is a strong plus * Bachelor's/Master's degree in Computer Science/ Engineering or a related field. ## Description Design and implement automated systems and pipelines for evaluating LLM outputs. * Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM-based evaluations * Collaborate with Engineering teams to create automated logic checks and validation tools. * Partner with Data Scientists to analyze evaluation results and optimize prompt and task structures. * Provide feedback loops to ensure evaluation guidelines align with LLM-based assessments. * Investigate how LLM-derived evaluations can enhance product reliability and user experience. * Recommend refinements to prompt engineering, evaluation strategies, and automation tools. * Stay informed on emerging trends in LLM evaluation, automated quality assessment, and AI toolchains. * Continuously improve and expand automated evaluation processes based on industry best practices. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Why LLMs Need Observability and How to Do It](https://www.wearedevelopers.com/videos/2117-why-llms-need-observability-and-how-to-do-it) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1520-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Give Your LLMs a Left Brain](https://www.wearedevelopers.com/videos/1160-give-your-llms-a-left-brain) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Prompt Engineering is a Job of the Past](https://www.wearedevelopers.com/magazine/342-prompt-engineering-is-a-job-of-the-past) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [The Prompt Engineer ✍️](https://www.wearedevelopers.com/magazine/216-the-prompt-engineer)