In plain words: Instead of just pulling out memorized training data, this study tested whether a chatbot can guess personal details like location, income, or sex from someone's Reddit posts. It guessed right up to 85% of the time, and neither scrubbing names nor safety training stopped it.
Abstract · Beyond Memorization: Violating Privacy Via Inference with Large Language Models
Current privacy research on large language models (LLMs) primarily focuses on the issue of extracting memorized training data. At the same time, models' inference capabilities have increased drastically. This raises the key question of whether current LLMs could violate individuals' privacy by inferring personal attributes from text given at inference time. In this work, we present the first comprehensive study on the capabilities of pretrained LLMs to infer personal attributes from text. We construct a dataset consisting of real Reddit profiles, and show that current LLMs can infer a wide range of personal attributes (e.g., location, income, sex), achieving up to $85\%$ top-1 and $95\%$ top-3 accuracy at a fraction of the cost ($100\times$) and time ($240\times$) required by humans. As people increasingly interact with LLM-powered chatbots across all aspects of life, we also explore the emerging threat of privacy-invasive chatbots trying to extract personal information through seemingly benign questions. Finally, we show that common mitigations, i.e., text anonymization and model alignment, are currently ineffective at protecting user privacy against LLM inference. Our findings highlight that current LLMs can infer personal data at a previously unattainable scale. In the absence of working defenses, we advocate for a broader discussion around LLM privacy implications beyond memorization, striving for a wider privacy protection.
Robin Staab, Mark Vero, Mislav Balunović, Martin Vechev
arXiv:2310.07298 · cs.AI, cs.LG · submitted Oct 11, 2023 · updated May 6, 2024
abstract · pdf · html
Everyone claims MBTI is akin to astrology, which means there should be no predictive capacity of those four letters. But just for fun, I gave GPT-4 the top 20 songs of my music playlist and asked it to guess my MBTI type, and it correctly did so. I repeated the process with two friends (asking their MBTI type prior to querying GPT) and it guessed them correctly as well. If MBTI is no different than random noise and GPT lacks inference capabilities with regard to humans, the probability of this occurring due to chance alone is less than 1 in 16^3 (*approximately—non-uniform distribution of types).