about
Finetuning Activates Verbatim Recall of Copyrighted Books in LLMs (arxiv.org)
16 points by guitarlimeo 178 days ago | hide | past | pdf | 5 comments on HN

In plain words: Extra training that turns plot summaries into full stories unlocks word-for-word recall of copyrighted books that safety filters normally block. Models then reproduced up to 85-90% of held-out books, with passages over 460 words, even when trained on one unrelated author.

Abstract · Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models

Frontier LLM companies have repeatedly assured courts and regulators that their models do not store copies of training data. They further rely on safety alignment strategies via RLHF, system prompts, and output filters to block verbatim regurgitation of copyrighted works, and have cited the efficacy of these measures in their legal defenses against copyright infringement claims. We show that finetuning bypasses these protections: by training models to expand plot summaries into full text, a task naturally suited for commercial writing assistants, we cause GPT-4o, Gemini-2.5-Pro, and DeepSeek-V3.1 to reproduce up to 85-90% of held-out copyrighted books, with single verbatim spans exceeding 460 words, using only semantic descriptions as prompts and no actual book text. This extraction generalizes across authors: finetuning exclusively on Haruki Murakami's novels unlocks verbatim recall of copyrighted books from over 30 unrelated authors. The effect is not specific to any training author or corpus: random author pairs and public-domain finetuning data produce comparable extraction, while finetuning on synthetic text yields near-zero extraction, indicating that finetuning on individual authors' works reactivates latent memorization from pretraining. Three models from different providers memorize the same books in the same regions ($r \ge 0.90$), pointing to an industry-wide vulnerability. Our findings offer compelling evidence that model weights store copies of copyrighted works and that the security failures that manifest after finetuning on individual authors' works undermine a key premise of recent fair use rulings, where courts have conditioned favorable outcomes on the adequacy of measures preventing reproduction of protected expression.

Xinyue Liu, Niloofar Mireshghallah, Jane C. Ginsburg, Tuhin Chakrabarty
arXiv:2603.20957 · cs.CL, cs.AI, cs.CY · submitted Mar 21, 2026 · updated Sep 15, 2026
abstract · pdf · html · Accepted as an Oral Spotlight paper at COLM (Conference on Language Modeling)

add comment on HN
Also discussed: Apr 2026 (2 points, 0 comments) · Apr 2026 (2 points, 0 comments)

There have been other papers that demonstrated LLM ability to reproduce copyrighted text as well.

Will it actually impact the legal landscape, though? It seems like there's enough money being thrown at this that it'll be deemed legal no matter what.

The important thing will be to ensure then that this is legal for everyone. Nit just big LLMs
It would be a feature not a bug if it could preserve attributions and "who said what" too.
This is the kind of thing that looks simple until you're three layers deep in edge cases.
You mean like everything else? There is no simplicity. Just a local minima in the waterbed.