about
Large-scale online deanonymization with LLMs (arxiv.org)
78 points by mellosouls 220 days ago | hide | past | pdf | 2 comments on HN

In plain words: An AI agent with internet access pulls identity clues from a person's pseudonymous posts, then searches for and double-checks who they are. It matched up to 68% of people while keeping 9 out of 10 matches correct, versus nearly none for non-AI methods.

Abstract

We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer participants at high precision, given pseudonymous online profiles and conversations alone, matching what would take hours for a dedicated human investigator. We then design attacks for the closed-world setting. Given two databases of pseudonymous individuals, each containing unstructured text written by or about that individual, we implement a scalable attack pipeline that uses LLMs to: (1) extract identity-relevant features, (2) search for candidate matches via semantic embeddings, and (3) reason over top candidates to verify matches and reduce false positives. Compared to classical deanonymization work (e.g., on the Netflix prize) that required structured data, our approach works directly on raw user content across arbitrary platforms. We construct three datasets with known ground-truth data to evaluate our attacks. The first links Hacker News to LinkedIn profiles, using cross-platform references that appear in the profiles. Our second dataset matches users across Reddit movie discussion communities; and the third splits a single user's Reddit history in time to create two pseudonymous profiles to be matched. In each setting, LLM-based methods substantially outperform classical baselines, achieving up to 68% recall at 90% precision compared to near 0% for the best non-LLM method. Our results show that the practical obscurity protecting pseudonymous users online no longer holds and that threat models for online privacy need to be reconsidered.

Simon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni, Nicholas Carlini, Florian Tramèr
arXiv:2602.16800 · cs.CR, cs.AI, cs.LG · submitted Feb 18, 2026 · updated Feb 25, 2026
abstract · pdf · html · 24 pages, 10 figures

add comment on HN
Also discussed: Feb 2026 (3 points, 2 comments) · Feb 2026 (1 point, 0 comments) · Feb 2026 (3 points, 0 comments)

Comments moved to https://news.ycombinator.com/item?id=47139716, which was posted by one of the authors. I hope that's ok!
I read long time ago opinion that anonymity and privacy are merely temporary side-products and rather short lived at that.

Piece argued that in medieval times and small cities people knew each other well and gossiped heavily (due to lack of other form of entertainment) and thus noone was truly anonymous. Privacy was taken away through prying eyes looking. According to it the only age of anonymity came before age of information which made it possible to crossreference various source of information which make it a blip on a human history timeline.

And today I'd say we're in the age of hyperinformation where enormous bodies of knowledge are compressed into (relatively) tiny LLMs which make crossreferencing even easier than before.