about
Evaluating faithfulness and content selection of LLMs in book-length summaries (arxiv.org)
71 points by passwordoops on Apr 9, 2024 | hide | past | pdf | 7 comments on HN

In plain words: People who had read each book checked 3,158 claims in AI-written summaries of 26 recent novels for accuracy and what got left out. AI judges failed to match the human verdicts, especially at spotting wrong claims, while one top model's summaries were most accurate.

Abstract · FABLES: Evaluating faithfulness and content selection in book-length summarization

While long-context large language models (LLMs) can technically summarize book-length documents (>100K tokens), the length and complexity of the documents have so far prohibited evaluations of input-dependent aspects like faithfulness. In this paper, we conduct the first large-scale human evaluation of faithfulness and content selection on LLM-generated summaries of fictional books. Our study mitigates the issue of data contamination by focusing on summaries of books published in 2023 or 2024, and we hire annotators who have fully read each book prior to the annotation task to minimize cost and cognitive burden. We collect FABLES, a dataset of annotations on 3,158 claims made in LLM-generated summaries of 26 books, at a cost of $5.2K USD, which allows us to rank LLM summarizers based on faithfulness: Claude-3-Opus significantly outperforms all closed-source LLMs, while the open-source Mixtral is on par with GPT-3.5-Turbo. An analysis of the annotations reveals that most unfaithful claims relate to events and character states, and they generally require indirect reasoning over the narrative to invalidate. While LLM-based auto-raters have proven reliable for factuality and coherence in other settings, we implement several LLM raters of faithfulness and find that none correlates strongly with human annotations, especially with regard to detecting unfaithful claims. Our experiments suggest that detecting unfaithful claims is an important future direction not only for summarization evaluation but also as a testbed for long-context understanding. Finally, we move beyond faithfulness by exploring content selection errors in book-length summarization: we develop a typology of omission errors related to crucial narrative elements and also identify a systematic over-emphasis on events occurring towards the end of the book.

Yekyung Kim, Yapei Chang, Marzena Karpinska, Aparna Garimella, Varun Manjunatha, Kyle Lo, Tanya Goyal, Mohit Iyyer
arXiv:2404.01261 · cs.CL, cs.AI · submitted Apr 1, 2024 · updated Sep 30, 2024
abstract · pdf · html · preprint - 39 pages

add comment on HN
Also discussed: Apr 2024 (2 points, 0 comments)

As far as I can tell, this study was almost entirely about fiction: https://github.com/mungg/FABLES/blob/main/booklist.md lists 26 books, only one of which is classified as non-fiction.

I would imagine that summaries of non-fiction books are evaluated quite differently from summaries of fiction.

I've been trying to figure out what prompts they used. The https://github.com/mungg/FABLES GitHub repo says this:

    Summary -- (str) Entire book summarized
    by one of five models: Mixtral,
    GPT-3.5-Turbo, GPT-4, GPT-4-Turbo, and
    Claude-3-Opus, using the hierarchical
    merging method described in Chang et al..
With a link to https://arxiv.org/pdf/2310.00785.pdf - which then links to another GitHub repository, https://github.com/lilakk/BooookScore which has a bunch of prompts in https://github.com/lilakk/BooookScore/tree/main/prompts

Which makes me think that this original paper isn't evaluating LLMs so much as it's evaluating that one particular prompting technique for long summaries.

Gemini Pro 1.5 has 1 million token context length, which should remove the need for weird hierarchical summary tricks. I wonder how well it would score?

The issue with non-fiction is that some information may come from the parametric memory, rather than source text supplied to the model. So then the issue is whether the model can process the text and summarize the book or is it cheating? Fiction published in the last few months is your best shot, though sure depending on the nature of non-fiction this evaluation can be different (but some things remain the same, like you want it to be faithful/factual and not omit important information)
Thanks, I hadn't caught that. Now I understand why the study used fiction rather than non-fiction.
No prob! there was not enough explanation in the paper about this (or none). As for the prompts used for merging, there is now a link in the github repo for this.
I didn't read the paper (just skimmed it), but no mention of Gemini 1.5 Pro? It is supposed to have the longest context window (1M available, 10M claimed in lab tests).
Let the potato rest a minute.

I suspect that they were preparing this for press by the time Gemini 1.5 Pro was released.

Gemini 1.5 Pro API was literary released yesterday. It took 11h+ for a person to evaluate a book (but that's if they sit down and do it all at once), it takes weeks to run something like this so yes... no gemini... yet...