about
Self-Taught Optimizer: Recursively self-improving code generation (arxiv.org)
3 points by SkyMarshal on Oct 6, 2023 | hide | past | pdf | 2 comments on HN

In plain words: A program that asks a language model to rewrite code so it scores better is used to rewrite itself, letting the improver upgrade its own instructions. On small tasks, the self-improved version wrote programs scoring significantly better than the original improver.

Abstract · Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation

Several recent advances in AI systems solve problems by providing a "scaffolding" program that structures multiple calls to language models (LMs) to generate better outputs. A scaffolding program is written in a programming language such as Python. In this work, we use a language-model-infused scaffolding program to improve itself. We start with a seed "improver" that improves an input program according to a given utility function by querying an LM several times and returning the best solution. We then run this seed improver to improve itself. Across a small set of downstream tasks, the resulting improved improver generates programs with significantly better performance than its seed improver. A variety of self-improvement strategies are proposed by the language model, including beam search, genetic algorithms, and simulated annealing. Since the language models themselves are not altered, this is not full recursive self-improvement. Nonetheless, it demonstrates that a modern language model, GPT-4 in our experiments, is capable of writing code that can call itself to improve itself. We consider concerns around the development of self-improving technologies and evaluate the frequency with which the generated code bypasses a sandbox.

Eric Zelikman, Eliana Lorch, Lester Mackey, Adam Tauman Kalai
arXiv:2310.02304 · cs.CL, cs.AI, cs.LG, stat.ML · submitted Oct 3, 2023 · updated Aug 16, 2024
abstract · pdf · Published as a conference paper at COLM 2024

add comment on HN
Also discussed: Oct 2023 (49 points, 10 comments) · Oct 2023 (1 point, 0 comments)

Hi, author here! We also have an accessible thread walking through this here: https://twitter.com/ericzelikman/status/1709721771937587541