about
Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best? (arxiv.org)
3 points by s1l3nt on Apr 8, 2025 | hide | past | pdf | 1 comment on HN

In plain words: It compares five automatic checkers that flag wrong answers from retrieval-based AI without needing a known correct answer, across six real uses. Across all six, a few checkers consistently caught errors, flagging wrong answers accurately and rarely missing them.

Abstract

This article surveys Evaluation models to automatically detect hallucinations in Retrieval-Augmented Generation (RAG), and presents a comprehensive benchmark of their performance across six RAG applications. Methods included in our study include: LLM-as-a-Judge, Prometheus, Lynx, the Hughes Hallucination Evaluation Model (HHEM), and the Trustworthy Language Model (TLM). These approaches are all reference-free, requiring no ground-truth answers/labels to catch incorrect LLM responses. Our study reveals that, across diverse RAG applications, some of these approaches consistently detect incorrect RAG responses with high precision/recall.

Ashish Sardana
arXiv:2503.21157 · cs.LG · submitted Mar 27, 2025 · updated Apr 7, 2025
abstract · pdf · 11 pages, 8 figures

add comment on HN

Is self-evaluation the blind spot of AI? Or can we trust LLMs to judge themselves? I benchmarked evaluation models/methods like LLM-as-a-judge, HHEM, Prometheus, Lynx, TLM across 6 RAG applications.

Evaluation models work surprisingly well in practice.

I hope continued research into reference-free evaluations helps users gain more confidence in their AI-generated outputs.