about
Bootstrapping Task Spaces for Self-Improvement (arxiv.org)
2 points by Anon84 on Sep 9, 2025 | hide | past | pdf | discuss on HN

In plain words: Instead of training on full revision chains of fixed length, it turns the most useful half-finished attempts into new starting tasks, so the agent practices one improvement step at a time. The resulting policies kept improving on problems past the revision depth seen in training.

Abstract

Progress in many task domains emerges from repeated revisions to previous solution attempts. Training agents that can reliably self-improve over such sequences at inference-time is a natural target for reinforcement learning (RL), yet the naive approach assumes a fixed maximum iteration depth, which can be both costly and arbitrary. We present Exploratory Iteration (ExIt), a family of autocurriculum RL methods that directly exploits the recurrent structure of self-improvement tasks to train LLMs to perform multi-step self-improvement at inference-time while only training on the most informative single-step iterations. ExIt grows a task space by selectively sampling the most informative intermediate, partial histories encountered during an episode for continued iteration, treating these starting points as new self-iteration task instances to train a self-improvement policy. ExIt can further pair with explicit exploration mechanisms to sustain greater task diversity. Across several domains, encompassing competition math, multi-turn tool-use, and machine learning engineering, we demonstrate that ExIt strategies, starting from either a single or many task instances, can produce policies exhibiting strong inference-time self-improvement on held-out task instances, and the ability to iterate towards higher performance over a step budget extending beyond the average iteration depth encountered during training.

Minqi Jiang, Andrei Lupu, Yoram Bachrach
arXiv:2509.04575 · cs.LG · submitted Sep 4, 2025 · updated Feb 22, 2026
abstract · pdf · html · TMLR, February 2026

add comment on HN