> Markdown version of [/jobs/ext/3011758-ai-evaluations-engineer-us-decision-intelligence](https://www.wearedevelopers.com/jobs/ext/3011758-ai-evaluations-engineer-us-decision-intelligence). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Evaluations Engineer, US Decision Intelligence - **Company:** Apple's Sales - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Continuous Integration, Distributed Systems, Graph Database, Apache Hadoop, Python (Programming Language), PostgreSQL, Machine Learning, RabbitMQ, Recommender Systems, Redis, Message Oriented Middleware, Software Engineering, SQL Databases, Large Language Models, Snowflake, Apache Spark, Caching, Information Technology, Data Analytics, Microservices - **Published:** September 20, 2026 - **Apply:** https://www.dice.com/job-detail/557960ed-1fdb-4e55-97c2-cb3d222459ab ## About the Role 5+ years of experience in data and AI-related fields such as AI engineering, software development, ML engineering, data science, or QA roles. Eagerness and ability to learn new skills and solve dynamic problems in an encouraging and expansive environment. Strong Python skills. Hands-on experience with AI evaluation techniques, such as Golden datasets, LLM-as-a-Judge, or rubric-based scoring. Experience with different LLM ecosystems (OpenAI, Anthropic, Gemini, etc.), RAG pipelines, vector databases (e.g., Pinecone, FAISS, Milvus, PostgreSQL). Proficiency in SQL and experience with at least one major data analytics platform, such as Hadoop, Spark, or Snowflake. Experience with CI/CD or release validation workflows. Experience working with data science teams on insights generation leveraging LLMs. Strong time management skills with the ability to collaborate across multiple teams. Able to balance competing priorities, long-term projects, and ad hoc requirements. Ability to work in a fast-paced, dynamic, constantly evolving business environment. Hands-on experience with Langfuse or similar tools for LLM observability. Comfortable working with product/domain experts to translate fuzzy correctness criteria into measurable rubrics or metrics. B.S. degree in Computer Science/Engineering, or equivalent work experience Preferred Qualifications Sound communication skills - expert at messaging domain and technical content, at a level appropriate for the audience. Strong ability to gain trust with stakeholders and senior leadership. Familiarity with embeddings, retrieval algorithms, agents, and data modeling for vector and graph databases. Other complementary technologies for distributed systems architecture and asynchronous messaging, agent communication, and caching like RabbitMQ, Redis, and Valkey are preferred. Experience working across global teams to ensure alignment of product development. Applied knowledge of GenAI and RAG strategies, microservices, recommendation systems, and context engineering. Working knowledge of agent evaluation concepts like trajectory vs. end-to-end vs. component-level evaluation, tool-call correctness. Advanced degree (MS or Ph.D.) in Economics, Electrical Engineering, Statistics, Data Science, or a similar quantitative field is preferred. ## Description We're seeking a visionary AI Evaluations Engineer to own the end-to-end evaluation pipeline for our AI products and agentic workflows. This role will focus on implementing and maintaining evaluation frameworks, instrumentation, and workflows that help us understand how well our AI systems perform, where they fail, and how they improve over time. You own the evaluation gate and the standards. This role will operate in both capacities, to augment existing AI roadmap, as well as innovate and trailblaze new frontier-technology projects, crafting AI experiences that reduce time to insight and catalyze decision making. ## Related Videos - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Beyond Kafka & RabbitMQ: Why NATS is the Future of Microservices Messaging](https://www.wearedevelopers.com/videos/1646-beyond-kafka-rabbitmq-why-nats-is-the-future-of-microservices-messaging) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) - [Developing ASP.NET Core Microservices with Dapr: A practical guide](https://www.wearedevelopers.com/videos/1528-developing-asp-net-core-microservices-with-dapr-a-practical-guide) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)