Evaluation
5 pieces on Evaluation, including 2 step-by-step guides.
Guides

Evaluate Multi-Turn Conversations with Ragas AspectCritic
Use Ragas AspectCritic to score multi-turn chatbot conversations on task completion, regulatory compliance, and brand voice with binary LLM judgements.
Read more →
Mastering Advanced RAG Evaluation: From Basic Metrics to LLM-as-a-Judge
Read more →
News & Analysis
ASR Benchmarks Are Gameable: 6 of 11 Top Models Reproduce Audio Errors
Hume AI tested 11 open-source ASR models and found six reproduce VoxPopuli's transcript errors even when audio contradicts them — exposing WER as a gameable metric.
Read more →

LLM Judges Carry Nine Measurable Biases: What to Do
DHS 2026 research catalogues nine exploitable biases in LLM-as-judge pipelines and shows grounded evaluators as the structural fix.
Read more →
Nine Measurable Biases That Corrupt LLM Judge Verdicts
A DHS 2026 workshop catalogued nine distinct biases in LLM-as-judge pipelines — and showed grounded evaluation as the structural fix.
Read more →