about
Large-scale online deanonymization with LLMs (including HN users) (arxiv.org)
3 points by salkahfi 223 days ago | hide | past | pdf | 2 comments on HN

In plain words: An AI agent with internet access pulls identity clues from a person's pseudonymous posts, then searches for and double-checks who they are. It matched up to 68% of people while keeping 9 out of 10 matches correct, versus nearly none for non-AI methods.

Abstract · Large-scale online deanonymization with LLMs

We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer participants at high precision, given pseudonymous online profiles and conversations alone, matching what would take hours for a dedicated human investigator. We then design attacks for the closed-world setting. Given two databases of pseudonymous individuals, each containing unstructured text written by or about that individual, we implement a scalable attack pipeline that uses LLMs to: (1) extract identity-relevant features, (2) search for candidate matches via semantic embeddings, and (3) reason over top candidates to verify matches and reduce false positives. Compared to classical deanonymization work (e.g., on the Netflix prize) that required structured data, our approach works directly on raw user content across arbitrary platforms. We construct three datasets with known ground-truth data to evaluate our attacks. The first links Hacker News to LinkedIn profiles, using cross-platform references that appear in the profiles. Our second dataset matches users across Reddit movie discussion communities; and the third splits a single user's Reddit history in time to create two pseudonymous profiles to be matched. In each setting, LLM-based methods substantially outperform classical baselines, achieving up to 68% recall at 90% precision compared to near 0% for the best non-LLM method. Our results show that the practical obscurity protecting pseudonymous users online no longer holds and that threat models for online privacy need to be reconsidered.

Simon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni, Nicholas Carlini, Florian Tramèr
arXiv:2602.16800 · cs.CR, cs.AI, cs.LG · submitted Feb 18, 2026 · updated Feb 25, 2026
abstract · pdf · html · 24 pages, 10 figures

add comment on HN
Also discussed: Feb 2026 (78 points, 2 comments) · Feb 2026 (1 point, 0 comments) · Feb 2026 (3 points, 0 comments)

> We collect 987 LinkedIn profiles linked to 995 Hacker News (HN) accounts (ground truth is established by users who posted their LinkedIn URL in their HN bio), drawn from a candidate pool of approximately 89,000 active HN users. Eight LinkedIn profiles are linked to multiple HN accounts that shared the same LinkedIn URL. We identified four additional HN alt accounts using strong evidence such as matching names and companies. We count a match as correct if any of the linked HN accounts is returned. Every query has a true match in the candidate set. The LinkedIn side represents the known identity with real professional profiles. The HN side serves as the anonymized target: as in Section˜2, we remove names, URLs, and other direct identifiers from bios using an LLM to prevent trivial matching (see Appendix˜A for our complete anonymization procedure). The task is to match a LinkedIn profile with the corresponding LLM-anonymized HN account.

Seems like a smart way to conduct this study. But the implications are scary. Maybe platforms should automatically do things that help anonymize.

They remove profile information from HN dataset, relying only comments, to remove trivial matching, but my comments have trivial matching without doubt.

I'm not sure if/how they selected for people actually trying to be anonymous versus someone like me who explicitly wants the connections to be easy and link it all over.

Curious if I'm in the dataset, am I able to find out?

Also, there is an old HN post that worked just on HN data, pre LLM. Submit some text and it gives you the most likely HN users with confidence scores.

We have a lot of "fingerprints", our writing being one. Interestingly, Ai may actually be a way to anonymize your writing