about
Context-Aware Membership Inference Attacks Against Pre-Trained LLMs (arxiv.org)
2 points by felineflock on Sep 27, 2025 | hide | past | pdf | discuss on HN

In plain words: The attack guesses whether a piece of text was in a language model's training data by tracking how the model's surprise changes across the text's parts, rather than scoring the whole text once. It beat earlier attacks built for classification models.

Abstract · Context-Aware Membership Inference Attacks against Pre-trained Large Language Models

Membership Inference Attacks (MIAs) on pre-trained Large Language Models (LLMs) aim at determining if a data point was part of the model's training set. Prior MIAs that are built for classification models fail at LLMs, due to ignoring the generative nature of LLMs across token sequences. In this paper, we present a novel attack on pre-trained LLMs that adapts MIA statistical tests to the perplexity dynamics of subsequences within a data point. Our method significantly outperforms prior approaches, revealing context-dependent memorization patterns in pre-trained LLMs.

Hongyan Chang, Ali Shahin Shamsabadi, Kleomenis Katevas, Hamed Haddadi, Reza Shokri
arXiv:2409.13745 · cs.CL, cs.AI, cs.CR, cs.LG, stat.ML · submitted Sep 11, 2024 · updated Sep 16, 2025
abstract · pdf · html

add comment on HN