about
No News Is Good News: A Critique of the One Billion Word Benchmark (arxiv.org)
15 points by CarrieLab on Nov 23, 2021 | hide | past | pdf | 1 comment on HN

In plain words: They tested a standard language-modeling test built from 2011 news by training models on web text from different years. Newer web text scored worse as news language drifted, and the test held harmful, outdated passages, making it a shaky measure of language skill.

Abstract · No News is Good News: A Critique of the One Billion Word Benchmark

The One Billion Word Benchmark is a dataset derived from the WMT 2011 News Crawl, commonly used to measure language modeling ability in natural language processing. We train models solely on Common Crawl web scrapes partitioned by year, and demonstrate that they perform worse on this task over time due to distributional shift. Analysis of this corpus reveals that it contains several examples of harmful text, as well as outdated references to current events. We suggest that the temporal nature of news and its distribution shift over time makes it poorly suited for measuring language modeling ability, and discuss potential impact and considerations for researchers building language models and evaluation datasets.

Helen Ngo, João G. M. Araújo, Jeffrey Hui, Nicholas Frosst
arXiv:2110.12609 · cs.CL, cs.LG · submitted Oct 25, 2021
abstract · pdf · html

add comment on HN

No Paper Is Good Paper: A Critique of Long Titles

The Arxiv One Billion Paper Benchmark was released in 2011, and is commonly used as a benchmark to writing academic papers. Analysis of this dataset shows that it contains several examples of sarcastic papers, as well as outdated references to current events, such as Support Vectors Machines. We suggest that the temporal nature of science makes this benchmark poorly suited to writing academic papers, and discuss potential impact and considerations for researchers building language models and evaluation datasets.

Conclusions

Papers written on top of other papers snap-shotted in time will display the inherent social bias and structural issues of that time. Therefore, people creating and using benchmarks, should realize that such a thing as drift exists, and we suggest they find ways around this. We encourage other paper writers to actively avoid using benchmarks where the training samples are always the same. This is a poor way to measure perplexity of language models and science. For better comparison, we suggest the training samples always change to reflect the current anti-bias Zeitgeist and that you cite our paper when doing so.