about
On the Risks of Stealing the Decoding Algorithms of Language Models (arxiv.org)
1 point by Irishsteve on Mar 27, 2023 | hide | past | pdf | discuss on HN

In plain words: By sending ordinary queries and studying the words that come back, an attacker can work out the rule that picks the next word from a model's probabilities, plus its secret settings. It worked on GPT-2, GPT-3 and GPT-Neo for $0.8 to $40.

Abstract · Stealing the Decoding Algorithms of Language Models

A key component of generating text from modern language models (LM) is the selection and tuning of decoding algorithms. These algorithms determine how to generate text from the internal probability distribution generated by the LM. The process of choosing a decoding algorithm and tuning its hyperparameters takes significant time, manual effort, and computation, and it also requires extensive human evaluation. Therefore, the identity and hyperparameters of such decoding algorithms are considered to be extremely valuable to their owners. In this work, we show, for the first time, that an adversary with typical API access to an LM can steal the type and hyperparameters of its decoding algorithms at very low monetary costs. Our attack is effective against popular LMs used in text generation APIs, including GPT-2, GPT-3 and GPT-Neo. We demonstrate the feasibility of stealing such information with only a few dollars, e.g., $\$0.8$, $\$1$, $\$4$, and $\$40$ for the four versions of GPT-3.

Ali Naseh, Kalpesh Krishna, Mohit Iyyer, Amir Houmansadr
arXiv:2303.04729 · cs.LG, cs.CL, cs.CR · submitted Mar 8, 2023 · updated Dec 1, 2023
abstract · pdf · html

add comment on HN