about
Generated Checklists Improve LLM Evaluation and Generation (arxiv.org)
22 points by hdvr on Oct 7, 2024 | hide | past | pdf | 1 comment on HN

In plain words: Instead of asking an AI judge for one overall ranking, this system turns each request into yes/no checks and grades the answer against each one. Its verdicts matched human preferences 52.2% of the time, up from 46.4% when the judge scored the answer directly.

Abstract · TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation

Given the widespread adoption and usage of Large Language Models (LLMs), it is crucial to have flexible and interpretable evaluations of their instruction-following ability. Preference judgments between model outputs have become the de facto evaluation standard, despite distilling complex, multi-faceted preferences into a single ranking. Furthermore, as human annotation is slow and costly, LLMs are increasingly used to make these judgments, at the expense of reliability and interpretability. In this work, we propose TICK (Targeted Instruct-evaluation with ChecKlists), a fully automated, interpretable evaluation protocol that structures evaluations with LLM-generated, instruction-specific checklists. We first show that, given an instruction, LLMs can reliably produce high-quality, tailored evaluation checklists that decompose the instruction into a series of YES/NO questions. Each question asks whether a candidate response meets a specific requirement of the instruction. We demonstrate that using TICK leads to a significant increase (46.4% $\to$ 52.2%) in the frequency of exact agreements between LLM judgements and human preferences, as compared to having an LLM directly score an output. We then show that STICK (Self-TICK) can be used to improve generation quality across multiple benchmarks via self-refinement and Best-of-N selection. STICK self-refinement on LiveBench reasoning tasks leads to an absolute gain of $+$7.8%, whilst Best-of-N selection with STICK attains $+$6.3% absolute improvement on the real-world instruction dataset, WildBench. In light of this, structured, multi-faceted self-improvement is shown to be a promising way to further advance LLM capabilities. Finally, by providing LLM-generated checklists to human evaluators tasked with directly scoring LLM responses to WildBench instructions, we notably increase inter-annotator agreement (0.194 $\to$ 0.256).

Jonathan Cook, Tim Rocktäschel, Jakob Foerster, Dennis Aumiller, Alex Wang
arXiv:2410.03608 · cs.AI, cs.CL, cs.HC, cs.LG · submitted Oct 4, 2024
abstract · pdf · html

add comment on HN

The authors propose a novel approach where checklists are automatically generated to systematically assess and guide LLM outputs, ensuring more comprehensive and reliable evaluations by LLMs. E.g. it increases in the frequency of exact agreements between LLM judgements and human preferences from 46.4% to 52.2%.

From my perspective it would be neat if the benchmarks would support more model types and not only the predominat GPTs, which only showed that they can relatively easy be scaled up, though it was never stated that they can model language better with the same resources (AFAIK).