about
Extracting Prompts by Inverting LLM Outputs (arxiv.org)
2 points by xcccube on May 27, 2024 | hide | past | pdf | 1 comment on HN

In plain words: A tool called output2prompt reads a model's answers and rebuilds the hidden prompt that produced them, using only ordinary questions—no peeking at internal scores and no trick questions. It recovered user and system prompts and carried over to models it never trained on.

Abstract

We consider the problem of language model inversion: given outputs of a language model, we seek to extract the prompt that generated these outputs. We develop a new black-box method, output2prompt, that learns to extract prompts without access to the model's logits and without adversarial or jailbreaking queries. In contrast to previous work, output2prompt only needs outputs of normal user queries. To improve memory efficiency, output2prompt employs a new sparse encoding techique. We measure the efficacy of output2prompt on a variety of user and system prompts and demonstrate zero-shot transferability across different LLMs.

Collin Zhang, John X. Morris, Vitaly Shmatikov
arXiv:2405.15012 · cs.CL, cs.LG · submitted May 23, 2024 · updated Oct 8, 2024
abstract · pdf · html

add comment on HN

We consider the problem of language model inversion: given outputs of a language model, we seek to extract the prompt that generated these outputs. We develop a new black-box method, output2prompt, that learns to extract prompts without access to the model's logits and without adversarial or jailbreaking queries. In contrast to previous work, output2prompt only needs outputs of normal user queries. To improve memory efficiency, output2prompt employs a new sparse encoding techique. We measure the efficacy of output2prompt on a variety of user and system prompts and demonstrate zero-shot transferability across different LLMs.