In plain words: Before each word is chosen, the system marks a random set of words 'green' and nudges the model toward them, leaving a hidden pattern. It barely hurt text quality and was detected from a short passage without the model's code or weights.
Abstract
Potential harms of large language models can be mitigated by watermarking model output, i.e., embedding signals into generated text that are invisible to humans but algorithmically detectable from a short span of tokens. We propose a watermarking framework for proprietary language models. The watermark can be embedded with negligible impact on text quality, and can be detected using an efficient open-source algorithm without access to the language model API or parameters. The watermark works by selecting a randomized set of "green" tokens before a word is generated, and then softly promoting use of green tokens during sampling. We propose a statistical test for detecting the watermark with interpretable p-values, and derive an information-theoretic framework for analyzing the sensitivity of the watermark. We test the watermark using a multi-billion parameter model from the Open Pretrained Transformer (OPT) family, and discuss robustness and security.
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein
arXiv:2301.10226 · cs.LG, cs.CL, cs.CR · submitted Jan 24, 2023 · updated May 1, 2024
abstract · pdf · html · 13 pages in the main body. Published at ICML 2023. Code is available at github.com/jwkirchenbauer/lm-watermarking
E.G. if it uses inter-textual spacing, then re-flow pagination would erode it. If it uses optional textual marking, a semantic read could replace ; instances with other constructs. If it does dependent word order for sentence end, or start or other stylometric changes, that too is probably statistically detectable.
To function for text fragments it will have to have some bitrate in the word stream, and therefore be subject to loss of sufficient bits to prevent reconstruction of the hash code (or whatever) it is based on.
I want this to exist, but I worry its a signal more than just the originator can detect, and therefore potentially can be defeated.
I very much hope they aren't proposing security by obscurity. If its public-private key based, then it will want the security community to review it with a fine tooth comb.
Oh dear. I see: Methods for keeping the watermark algorithm secret but available via API are discussed in Section 5.
So they propose a large set of tokens, and propose a 50/50 style selection of which have meaning and which don't to increase the surface of cost to detect the tokens and "flip" them, therefore permitting them to argue the watermark may have been found, but could not be entirely eroded.
In an earlier HN thread I was (incorrectly) accused of using a GPT to write my responses. Shortly after, I read of others, and a few days after that Scott Aarenson talked about wanting to deploy a watermark/signature method. I do not see myself as a trigger in that I am just sui generis an instance of why he said they need to: It turns out the simplistic "you are a bot" people have a horrendous false positive rate, and a lot of people who write like me are at risk of being labelled bots, when we aren't.
(this whine was not produced with ChatGPT)