about
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness (arxiv.org)
1 point by alastairr on Feb 23, 2024 | hide | past | pdf | discuss on HN

In plain words: TrustScore checks whether a chatbot's answer matches what the same model knows elsewhere, spotting shaky answers without needing a correct answer to compare against. It agreed with human judgments better than other no-reference checks, matching the ones that use reference answers.

Abstract

Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, prompting a surge in their practical applications. However, concerns have arisen regarding the trustworthiness of LLMs outputs, particularly in closed-book question-answering tasks, where non-experts may struggle to identify inaccuracies due to the absence of contextual or ground truth information. This paper introduces TrustScore, a framework based on the concept of Behavioral Consistency, which evaluates whether an LLMs response aligns with its intrinsic knowledge. Additionally, TrustScore can seamlessly integrate with fact-checking methods, which assesses alignment with external knowledge sources. The experimental results show that TrustScore achieves strong correlations with human judgments, surpassing existing reference-free metrics, and achieving results on par with reference-based metrics.

Danna Zheng, Danyang Liu, Mirella Lapata, Jeff Z. Pan
arXiv:2402.12545 · cs.CL · submitted Feb 19, 2024 · updated May 6, 2024
abstract · pdf · html

add comment on HN