In plain words: It compares five automatic checkers that flag wrong answers from retrieval-based AI without needing a known correct answer, across six real uses. Across all six, a few checkers consistently caught errors, flagging wrong answers accurately and rarely missing them.
Abstract
This article surveys Evaluation models to automatically detect hallucinations in Retrieval-Augmented Generation (RAG), and presents a comprehensive benchmark of their performance across six RAG applications. Methods included in our study include: LLM-as-a-Judge, Prometheus, Lynx, the Hughes Hallucination Evaluation Model (HHEM), and the Trustworthy Language Model (TLM). These approaches are all reference-free, requiring no ground-truth answers/labels to catch incorrect LLM responses. Our study reveals that, across diverse RAG applications, some of these approaches consistently detect incorrect RAG responses with high precision/recall.
Ashish Sardana
arXiv:2503.21157 · cs.LG · submitted Mar 27, 2025 · updated Apr 7, 2025
abstract · pdf · 11 pages, 8 figures
Evaluation models work surprisingly well in practice.
I hope continued research into reference-free evaluations helps users gain more confidence in their AI-generated outputs.