about
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Models (arxiv.org)
47 points by Jimmc414 on Apr 30, 2024 | hide | past | pdf | 4 comments on HN

In plain words: Instead of one big AI model grading another's answers, this uses a panel of smaller models from different families and combines their scores. The panel graded more accurately than the single big judge, favored its own family less, and cost over seven times less.

Abstract · Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

As Large Language Models (LLMs) have become more advanced, they have outpaced our abilities to accurately evaluate their quality. Not only is finding data to adequately probe particular model properties difficult, but evaluating the correctness of a model's freeform generation alone is a challenge. To address this, many evaluations now rely on using LLMs themselves as judges to score the quality of outputs from other LLMs. Evaluations most commonly use a single large model like GPT4. While this method has grown in popularity, it is costly, has been shown to introduce intramodel bias, and in this work, we find that very large models are often unnecessary. We propose instead to evaluate models using a Panel of LLm evaluators (PoLL). Across three distinct judge settings and spanning six different datasets, we find that using a PoLL composed of a larger number of smaller models outperforms a single large judge, exhibits less intra-model bias due to its composition of disjoint model families, and does so while being over seven times less expensive.

Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, Patrick Lewis
arXiv:2404.18796 · cs.CL, cs.AI · submitted Apr 29, 2024 · updated May 1, 2024
abstract · pdf · html

add comment on HN

This reminds me of science fiction author Peter Watts' novella "The Freeze-Frame Revolution", where a space ship has two AIs: one that has been running for a million years and another that is reboot daily to start with a fresh state. The long-running AI confers with reboot AI for a second opinion on important issues. The second AI doesn't know it's "killed" daily, but eventually starts to suspect. And this is just a small subplot! If you like hard SF jam-packed with big ideas, I highly recommend "The Freeze-Frame Revolution" and Watts' other novels.
I'll have to read it!

I've been mulling over a sci-fi setting focused on keeping humans relevant in the context of advanced technology. One of the core ideas is that AIs decay without human interaction. State backups don't prevent it because noticing clues they've been out just makes it worse.

I guess the whole "it's easier to be a critic than a writer" thing applies to LLMs too.
I was going to knock this for incrementally improving performance while massively increasing costs, but it's actually 7x less expensive. Not bad.