about
LLMmap: Fingerprinting for Large Language Models (arxiv.org)
2 points by emk_709 on Oct 30, 2024 | hide | past | pdf | 2 comments on HN

In plain words: LLMmap sends a few carefully chosen questions to an app and reads the answers to work out which language model version is running behind it. Just 8 questions identified 42 versions with over 95% accuracy, even with hidden instructions or extra processing layers.

Abstract · LLMmap: Fingerprinting For Large Language Models

We introduce LLMmap, a first-generation fingerprinting technique targeted at LLM-integrated applications. LLMmap employs an active fingerprinting approach, sending carefully crafted queries to the application and analyzing the responses to identify the specific LLM version in use. Our query selection is informed by domain expertise on how LLMs generate uniquely identifiable responses to thematically varied prompts. With as few as 8 interactions, LLMmap can accurately identify 42 different LLM versions with over 95% accuracy. More importantly, LLMmap is designed to be robust across different application layers, allowing it to identify LLM versions--whether open-source or proprietary--from various vendors, operating under various unknown system prompts, stochastic sampling hyperparameters, and even complex generation frameworks such as RAG or Chain-of-Thought. We discuss potential mitigations and demonstrate that, against resourceful adversaries, effective countermeasures may be challenging or even unrealizable.

Dario Pasquini, Evgenios M. Kornaropoulos, Giuseppe Ateniese
arXiv:2407.15847 · cs.CR, cs.AI · submitted Jul 22, 2024 · updated Feb 10, 2025
abstract · pdf · html · Appearing in the proceedings of the 34th USENIX Security Symposium

add comment on HN

"Social engineering" for modern era.

Thanks for posting the paper

Haha! I haven't thought about it this way, but you are correct.