about
Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (arxiv.org)
2 points by mnk47 on Nov 22, 2024 | hide | past | pdf | discuss on HN

In plain words: Treats benchmark questions as a random sample from a huge pool of possible questions, so standard statistics can put error bars on scores and size tests properly. Unlike today's single-number scores, this shows how to tell real model differences from luck and cut noise.

Abstract

Evaluations are critical for understanding the capabilities of large language models (LLMs). Fundamentally, evaluations are experiments; but the literature on evaluations has largely ignored the literature from other sciences on experiment analysis and planning. This article shows researchers with some training in statistics how to think about and analyze data from language model evaluations. Conceptualizing evaluation questions as having been drawn from an unseen super-population, we present formulas for analyzing evaluation data, measuring differences between two models, and planning an evaluation experiment. We make a number of specific recommendations for running language model evaluations and reporting experiment results in a way that minimizes statistical noise and maximizes informativeness.

Evan Miller
arXiv:2411.00640 · stat.AP, cs.CL · submitted Nov 1, 2024
abstract · pdf · html · 14 pages

add comment on HN
Also discussed: Aug 2026 (3 points, 0 comments) · Nov 2024 (1 point, 0 comments)