about
An Archaeology of Books Known to ChatGPT/GPT-4 (arxiv.org)
3 points by belter on May 4, 2023 | hide | past | pdf | 1 comment on HN

In plain words: By hiding a book's name and asking the model to fill in the blank, they test which books ChatGPT and GPT-4 have memorized. The models knew many copyrighted books, especially ones common online, and did better on those than on books they hadn't memorized.

Abstract · Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4

In this work, we carry out a data archaeology to infer books that are known to ChatGPT and GPT-4 using a name cloze membership inference query. We find that OpenAI models have memorized a wide collection of copyrighted materials, and that the degree of memorization is tied to the frequency with which passages of those books appear on the web. The ability of these models to memorize an unknown set of books complicates assessments of measurement validity for cultural analytics by contaminating test data; we show that models perform much better on memorized books than on non-memorized books for downstream tasks. We argue that this supports a case for open models whose training data is known.

Kent K. Chang, Mackenzie Cramer, Sandeep Soni, David Bamman
arXiv:2305.00118 · cs.CL · submitted Apr 28, 2023 · updated Oct 20, 2023
abstract · pdf · html · EMNLP 2023 camera-ready (16 pages, 4 figures)

add comment on HN