about
Lessons from the trenches on reproducible evaluation of language models (arxiv.org)
42 points by veryluckyxyz on May 25, 2024 | hide | past | pdf | 3 comments on HN

In plain words: Drawing on three years of building a widely used language-model testing tool, this report collects the pitfalls that make scores unreliable and hard to compare. Small setup choices can change results, so it urges teams to fix and report details most evaluations leave out.

Abstract · Lessons from the Trenches on Reproducible Evaluation of Language Models

Reliable evaluation of language models (LMs) remains an open challenge. Re- searchers and engineers face methodological issues such as the sensitivity of models to evaluation setup, difficulty of proper comparisons across methods, and the lack of reproducibility and transparency. Evaluation difficulties are exacer- bated by the fracturing and siloing of information about conventions and common practices. In this paper we draw on three years of experience in evaluating large lan- guage models (LMs) as developers of the popular Language Model Evaluation Harness (lm-eval) (Gao et al., 2023) framework to provide guidance and lessons for the field moving forward. We document a variety of challenges faced by prac- titioners and provide concrete instances where these challenges or the absence of best practices have come into effect. We make recommendations to the field for improving evaluation rigor and confidence, and attempt to codify much of the tacit or folk knowledge surrounding LM evaluation, for a solid ground to move forward.

Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, et al.
arXiv:2405.14782 · cs.CL · submitted May 23, 2024 · updated May 31, 2026
abstract · pdf · html

add comment on HN
Also discussed: May 2024 (1 point, 0 comments)

One point they don’t seem to spend much time on is also the difficulty in reproducing outputs in closed-source models. Setting temperature to 0 and setting seeds doesn’t always seem to be enough to get exactly the same results for a given prompt
Are there other parameters that affect the output?
Library versions with slightly different numeric rounding errors, alternative implementation, runtimes, and hardware variation could all lead to reproduction challenges.

There is no obligation for tf’s sigmoid implementation to exactly match PyTorch’s output. The same is true for nvidia vs amd.