Topic mix

LLM evaluation

13 moments from 12 videos · 45:08 min total

Assess the quality and accuracy of your language model outputs. These curated talk segments cover automated benchmarking, human-in-the-loop workflows, and testing metrics.

Evals vs. Evil - AI and Package Security - Laurie Voss
Play section Defining and implementing LLM evaluation strategies
Defining and implementing LLM evaluation strategies thumbnail

Defining and implementing LLM evaluation strategies

Because traditional unit tests fail on variable LLM outputs, utilizing a separate model for evaluation serves as an effective testing substitute.

Lessons Learned Building a GenAI Powered App
Play section Evaluating LLM response accuracy using secondary LLMs
Evaluating LLM response accuracy using secondary LLMs thumbnail

Evaluating LLM response accuracy using secondary LLMs

Creating an automated validation loop where a secondary model grades the initial facts.

LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices
Play section Executing fine-tuning and LLM evaluation API workflows
Executing fine-tuning and LLM evaluation API workflows thumbnail

Executing fine-tuning and LLM evaluation API workflows

Pulling foundation models and datasets via API requests enables automated container training and adapter benchmarking.

AI as a Test Designer: Transforming Experience into Automated Testing
Play section Validating test cases using the LLM judge concept
Validating test cases using the LLM judge concept thumbnail

Validating test cases using the LLM judge concept

Deploying a secondary LLM model evaluates newly generated test scenarios for quality and relevance against baseline usage patterns.

Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source
Play section Addressing system transparency when utilizing an LLM judge
Addressing system transparency when utilizing an LLM judge thumbnail

Addressing system transparency when utilizing an LLM judge

Leveraging deep evaluation traces clarifies the logical steps and exact reasoning behind individual automated assessment scores.

Play section Programmatic model evaluation and custom metrics via MLflow
Programmatic model evaluation and custom metrics via MLflow thumbnail

Programmatic model evaluation and custom metrics via MLflow

Using an LLM as a judge alongside clear guidelines within MLflow creates an automated baseline for continuous improvement.

Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems
Play section Introducing LLMs as judges for automated testing
Introducing LLMs as judges for automated testing thumbnail

Introducing LLMs as judges for automated testing

Using independent language models to critically review chatbot outputs at scale for correctness and consistency.

MAD About Software Design - When AI Architects Argue
Play section Evaluating AI debate quality against formal architecture katas
Evaluating AI debate quality against formal architecture katas thumbnail

Evaluating AI debate quality against formal architecture katas

Systematic testing reveals that clarifications improve results while summarization and excessive debate rounds degrade design quality.

Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models
Play section Evaluating prompt accuracy against expected test outcomes
Evaluating prompt accuracy against expected test outcomes thumbnail

Evaluating prompt accuracy against expected test outcomes

Establishing test datasets and using LLM judges to ensure accuracy and coherence.

Give Your LLMs a Left Brain
Play section Challenges of adopting LLMs for enterprise applications
Challenges of adopting LLMs for enterprise applications thumbnail

Challenges of adopting LLMs for enterprise applications

Commercial LLMs trained on broad datasets face significant obstacles when applied to specific enterprise contexts and problems.

Stop Guessing, Start Measuring: Evaluating RAG Systems with Synthetic Test Data
Play section Executing automated evaluation suites leveraging LLMs as judges
Executing automated evaluation suites leveraging LLMs as judges thumbnail

Executing automated evaluation suites leveraging LLMs as judges

Scoring synthetic dataset queries programmatically against system responses provides an objective baseline for fine-tuning configuration changes.

11 Principles for Evaluating AI Dev Tools
Play section Separating autonomous code generation from verification concerns
Separating autonomous code generation from verification concerns thumbnail

Separating autonomous code generation from verification concerns

Reducing cognitive bias in LLM outputs requires independent evaluation harnesses rather than letting agents review themselves.

Hack Me If You Can: Designing Unbreakable LLM Guardrails
Play section Implementing regex, classifiers, and LLM-as-a-judge guardrails
Implementing regex, classifiers, and LLM-as-a-judge guardrails thumbnail

Implementing regex, classifiers, and LLM-as-a-judge guardrails

Teams construct layered defenses by mixing cheap regex rules with lightweight structural classifiers and analytical evaluation models.

Your mix. Instantly.

More mixes