about
Language Model Cascades (arxiv.org)
1 point by 2517AD on Jul 23, 2022 | hide | past | pdf | discuss on HN

In plain words: Chaining prompts, checking answers, or calling tools can be written as one program where a model's answers are random choices and the code decides what happens next. This shows tricks like chain-of-thought, answer checkers, and tool use are all special cases of one framework.

Abstract

Prompted models have demonstrated impressive few-shot learning abilities. Repeated interactions at test-time with a single model, or the composition of multiple models together, further expands capabilities. These compositions are probabilistic models, and may be expressed in the language of graphical models with random variables whose values are complex data types such as strings. Cases with control flow and dynamic structure require techniques from probabilistic programming, which allow implementing disparate model structures and inference strategies in a unified language. We formalize several existing techniques from this perspective, including scratchpads / chain of thought, verifiers, STaR, selection-inference, and tool use. We refer to the resulting programs as language model cascades.

David Dohan, Winnie Xu, Aitor Lewkowycz, Jacob Austin, David Bieber, Raphael Gontijo Lopes, Yuhuai Wu, Henryk Michalewski, Rif A. Saurous, Jascha Sohl-dickstein, Kevin Murphy, Charles Sutton
arXiv:2207.10342 · cs.CL, cs.AI · submitted Jul 21, 2022 · updated Jul 28, 2022
abstract · pdf · html · Presented as spotlight at the Beyond Bases workshop at ICML 2022 (https://beyond-bayes.github.io)

add comment on HN