about
LLMs can hide arbitrary information in their responses, undetectably (arxiv.org)
1 point by cryptohell on Jan 22, 2024 | hide | past | pdf | discuss on HN

In plain words: A secret message is slipped into a chatbot's reply so that only someone holding a special key can pull it back out. Without the key, the reply is provably impossible to tell apart from a normal one, and the text quality stays the same.

Abstract · Excuse me, sir? Your language model is leaking (information)

We introduce a cryptographic method to hide an arbitrary secret payload in the response of a Large Language Model (LLM). A secret key is required to extract the payload from the model's response, and without the key it is provably impossible to distinguish between the responses of the original LLM and the LLM that hides a payload. In particular, the quality of generated text is not affected by the payload. Our approach extends a recent result of Christ, Gunn and Zamir (2023) who introduced an undetectable watermarking scheme for LLMs.

Or Zamir
arXiv:2401.10360 · cs.CR, cs.LG · submitted Jan 18, 2024
abstract · pdf · html

add comment on HN