about
Long-form factuality in large language models (arxiv.org)
18 points by rootforce on Mar 29, 2024 | hide | past | pdf | 3 comments on HN

In plain words: A checker that splits a long AI answer into single facts and looks each one up online to see if search results back it, tested on thousands of open-ended questions. It matched human fact-checkers 72% of the time while costing far less.

Abstract

Large language models (LLMs) often generate content that contains factual errors when responding to fact-seeking prompts on open-ended topics. To benchmark a model's long-form factuality in open domains, we first use GPT-4 to generate LongFact, a prompt set comprising thousands of questions spanning 38 topics. We then propose that LLM agents can be used as automated evaluators for long-form factuality through a method which we call Search-Augmented Factuality Evaluator (SAFE). SAFE utilizes an LLM to break down a long-form response into a set of individual facts and to evaluate the accuracy of each fact using a multi-step reasoning process comprising sending search queries to Google Search and determining whether a fact is supported by the search results. Furthermore, we propose extending F1 score as an aggregated metric for long-form factuality. To do so, we balance the percentage of supported facts in a response (precision) with the percentage of provided facts relative to a hyperparameter representing a user's preferred response length (recall). Empirically, we demonstrate that LLM agents can outperform crowdsourced human annotators - on a set of ~16k individual facts, SAFE agrees with crowdsourced human annotators 72% of the time, and on a random subset of 100 disagreement cases, SAFE wins 76% of the time. At the same time, SAFE is more than 20 times cheaper than human annotators. We also benchmark thirteen language models on LongFact across four model families (Gemini, GPT, Claude, and PaLM-2), finding that larger language models generally achieve better long-form factuality. LongFact, SAFE, and all experimental code are available at https://github.com/google-deepmind/long-form-factuality.

Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Jie Huang, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, Quoc V. Le
arXiv:2403.18802 · cs.CL, cs.AI, cs.LG · submitted Mar 27, 2024 · updated Nov 7, 2024
abstract · pdf · html · NeurIPS 2024; 72 pages, 18 figures, 30 tables. Code at https://github.com/google-deepmind/long-form-factuality

add comment on HN
Also discussed: Apr 2024 (26 points, 16 comments) · Apr 2024 (2 points, 0 comments) · Mar 2024 (1 point, 0 comments) · Mar 2024 (1 point, 0 comments)

Hmmm... checking against external sources is an interesting idea -- but using Google as a source of ground truth is a little bit tricky, given how often these days Google itself is spitting up confabulated AI-generated crud (or other low-quality stuff).
Use books and papers from the Library of Genesis - that gives you good context, even while the search engines collapse

End Google

Long live The Library!

For those interested in using search-augmented "reasoning", I implemented something similar in Emerging Trajectories[1], an open source package that forecasts geopolitical and economic events. We extract facts[2] from various websites (Google searches, news articles, RSS feeds) and have the LLM generate a hypothesis on a metric.

We're tracking the info forecasts to see how well this does for future events. For example, we're pitting the LLMs against each other to predict March 2024 CPI[3].

[1] https://emergingtrajectories.com/

[2] Sample code: https://github.com/wgryc/emerging-trajectories/blob/main/eme...

[3] https://emergingtrajectories.com/a/statement/28