about
Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling (arxiv.org)
1 point by ftxbro on May 24, 2023 | hide | past | pdf | discuss on HN

In plain words: Instead of a person writing step-by-step solution instructions, this system repeatedly writes new ones using the best found so far as examples, keeping those that solve practice problems well. Across 20 reasoning tasks it beat human-written instructions by 9.4 points on average.

Abstract

We introduce Reprompting, an iterative sampling algorithm that automatically learns the Chain-of-Thought (CoT) recipes for a given task without human intervention. Through Gibbs sampling, Reprompting infers the CoT recipes that work consistently well for a set of training samples by iteratively sampling new recipes using previously sampled recipes as parent prompts to solve other training problems. We conduct extensive experiments on 20 challenging reasoning tasks. Results show that Reprompting outperforms human-written CoT prompts substantially by +9.4 points on average. It also achieves consistently better performance than the state-of-the-art prompt optimization and decoding algorithms.

Weijia Xu, Andrzej Banburski-Fahey, Nebojsa Jojic
arXiv:2305.09993 · cs.LG, cs.AI, cs.CL · submitted May 17, 2023 · updated May 23, 2024
abstract · pdf · html · ICML 2024

add comment on HN