about
Do the simplest LLM prompts work best? (arxiv.org)
1 point by MichaelMoser123 on May 13, 2023 | hide | past | pdf | 1 comment on HN

In plain words: Prompts written in language a model finds easy to predict work better, so the study measures each prompt's predictability and picks the most familiar ones. Expanding a few hand-written prompts by rewording them and choosing the most predictable gave significant task gains.

Abstract · Demystifying Prompts in Language Models via Perplexity Estimation

Language models can be prompted to perform a wide variety of zero- and few-shot learning problems. However, performance varies significantly with the choice of prompt, and we do not yet understand why this happens or how to pick the best prompts. In this work, we analyze the factors that contribute to this variance and establish a new empirical hypothesis: the performance of a prompt is coupled with the extent to which the model is familiar with the language it contains. Over a wide range of tasks, we show that the lower the perplexity of the prompt is, the better the prompt is able to perform the task. As a result, we devise a method for creating prompts: (1) automatically extend a small seed set of manually written prompts by paraphrasing using GPT3 and backtranslation and (2) choose the lowest perplexity prompts to get significant gains in performance.

Hila Gonen, Srini Iyer, Terra Blevins, Noah A. Smith, Luke Zettlemoyer
arXiv:2212.04037 · cs.CL · submitted Dec 8, 2022 · updated Sep 12, 2024
abstract · pdf · html · Published in Findings of EMNLP 2023

add comment on HN

from the article: "we devise the following straightforward procedure:

1. Obtain a small set of manually created prompts for the task.

2. Expand the set of prompts with automatic paraphrasing using a LM (e.g., GPT3) and backtranslation (see Section 3).

3. Rank the list of prompts by perplexity (aver- aged on a representative sample of task inputs, e.g. 1,000).

4. Choose the k (e.g., 3) lowest perplexity prompts.

Using this algorithm, we show empirically that it is best to prioritize experimenting with the lowest perplexity prompts, as they perform better than manual prompts on average, and are more stable"