about
LLMs can hide text in other text of the same length (arxiv.org)
2 points by goplayoutside 340 days ago | hide | past | pdf | 1 comment on HN

In plain words: A trick lets a language model bury a secret message inside a different but sensible text of the same length, then pull it out. Even a small model can hide and recover a message as long as this summary on a laptop in seconds.

Abstract

A meaningful text can be hidden inside another, completely different yet still coherent and plausible, text of the same length. For example, a tweet containing a harsh political critique could be embedded in a tweet that celebrates the same political leader, or an ordinary product review could conceal a secret manuscript. This uncanny state of affairs is now possible thanks to Large Language Models, and in this paper we present Calgacus, a simple and efficient protocol to achieve it. We show that even modest 8-billion-parameter open-source LLMs are sufficient to obtain high-quality results, and a message as long as this abstract can be encoded and decoded locally on a laptop in seconds. The existence of such a protocol demonstrates a radical decoupling of text from authorial intent, further eroding trust in written communication, already shaken by the rise of LLM chatbots. We illustrate this with a concrete scenario: a company could covertly deploy an unfiltered LLM by encoding its answers within the compliant responses of a safe model. This possibility raises urgent questions for AI safety and challenges our understanding of what it means for a Large Language Model to know something.

Antonio Norelli, Michael Bronstein
arXiv:2510.20075 · cs.AI, cs.CL, cs.CR, cs.LG · submitted Oct 22, 2025 · updated Jan 16, 2026
abstract · pdf · html · 21 pages, main paper 9 pages. v5 contains an Italian translation of this paper by the author

add comment on HN
Also discussed: Jul 2026 (5 points, 0 comments) · May 2026 (5 points, 0 comments)

https://x.com/rohanpaul_ai/status/1982222641345057263

>The paper shows how an LLM can hide a full message inside another text of equal length.

>It runs in seconds on a laptop with 8B open models.

>First, pass the secret through an LLM and record, for each token, the rank of the actual next token.

>Then prompt the model to write on a chosen topic, and force it to pick tokens at those ranks.

>The result reads normally on that topic and has the same token count as the secret.

>With the same model and prompt, anyone can reverse the steps and recover the exact original.

>These covers look natural to people, but models usually rate them less likely than the originals.

>Quality is best when the model predicts the hidden text well, and worse for unusual domains or weaker models.

>Security comes from the secret prompt and the exact model, and it gives the sender believable deniability.

>One risk is hiding harmful answers inside safe replies for later extraction by a local model.