In plain words: Instead of trusting Wikipedia, books, and news as the gold standard, they tested the filter that picks web text for GPT-3 using U.S. high school newspapers. It favored writing from bigger, wealthier, urban schools, and its quality scores did not match factuality or literary acclaim.
Abstract · Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection
Language models increasingly rely on massive web dumps for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and newswire often serve as anchors for automatically selecting web text most suitable for language modeling, a process typically referred to as quality filtering. Using a new dataset of U.S. high school newspaper articles -- written by students from across the country -- we investigate whose language is preferred by the quality filter used for GPT-3. We find that newspapers from larger schools, located in wealthier, educated, and urban ZIP codes are more likely to be classified as high quality. We then demonstrate that the filter's measurement of quality is unaligned with other sensible metrics, such as factuality or literary acclaim. We argue that privileging any corpus as high quality entails a language ideology, and more care is needed to construct training corpora for language models, with better transparency and justification for the inclusion or exclusion of various texts.
Suchin Gururangan, Dallas Card, Sarah K. Dreier, Emily K. Gade, Leroy Z. Wang, Zeyu Wang, Luke Zettlemoyer, Noah A. Smith
arXiv:2201.10474 · cs.CL, cs.AI · submitted Jan 25, 2022 · updated Jan 26, 2022
abstract · pdf · html
Their argument seems predicated on the idea that either the author is the only writer and the text leapt from his head like Athena fully formed. (like these comments of mine, surely), or their entire sample set of student newspapers all had equally competent sub-editors. I'd say their argument that privileging any text is a "language ideology" is weak because the percieved quality of the writing should be attributed to the additional work that went into its editing, whereas they're saying it's due to the authors social status based on zip code. Chances are, the smaller school papers are just some yahoo publishing their own copy.
Too many holes. It seems to just elevate the same critical theory as a pretext for asserting a qualification to govern GPT model training, by using the same problematizations it uses on everything else. (e.g. call it x'ist until you control it, invent an unsolvable problem only you can manage, dilute and destabilize consensus with exogenous concerns, etc.) I'd agree a lot of good stuff is probably not making it into language models because it's not edited (or ironically, not gatekept), but I'm not sure the authors are really sincere about improving language models. To me, they're using a very narrow interpretation of quality writing to assert that GPT models require governance and political accountability.