about
We Should Separate Memorization from Copyright (arxiv.org)
1 point by 50kIters 236 days ago | hide | past | pdf | discuss on HN

In plain words: Technical tests that try to rebuild training data measure memorization, not illegal copying, so they can't tell whether a model broke copyright law. The paper separates signals that suggest infringement from those showing ordinary learning, and proposes judging each output by its legal risk.

Abstract

The widespread use of foundation models has introduced a new risk factor of copyright issue. This issue is leading to an active, lively and on-going debate amongst the data-science community as well as amongst legal scholars. Where claims and results across both sides are often interpreted in different ways and leading to different implications. Our position is that much of the technical literature relies on traditional reconstruction techniques that are not designed for copyright analysis. As a result, memorization and copying have been conflated across both technical and legal communities and in multiple contexts. We argue that memorization, as commonly studied in data science, should not be equated with copying and should not be used as a proxy for copyright infringement. We distinguish technical signals that meaningfully indicate infringement risk from those that instead reflect lawful generalization or high-frequency content. Based on this analysis, we advocate for an output-level, risk-based evaluation process that aligns technical assessments with established copyright standards and provides a more principled foundation for research, auditing, and policy.

Adi Haviv, Niva Elkin-Koren, Uri Hacohen, Roi Livni, Shay Moran
arXiv:2602.08632 · cs.CY, cs.AI, cs.CL, cs.CV, cs.LG · submitted Feb 9, 2026
abstract · pdf · html

add comment on HN