In plain words: A learned curator picks which practice problems to feed a language model next, choosing the ones it predicts will improve the model most and updating as the model changes. It beat uniform sampling and other curricula, raising scores 28.6% on a hard math contest.
Abstract · Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training
Post-training large foundation models with reinforcement learning typically relies on massive and heterogeneous datasets, making effective curriculum learning both critical and challenging. In this work, we propose ACTOR-CURATOR, a scalable and fully automated curriculum learning framework for reinforcement learning post-training of large language models (LLMs). ACTOR-CURATOR learns a neural curator that dynamically selects training problems from large problem banks by directly optimizing for expected policy performance improvement. We formulate problem selection as a non-stationary stochastic bandit problem, derive a principled loss function based on online stochastic mirror descent, and establish regret guarantees under partial feedback. Empirically, ACTOR-CURATOR consistently outperforms uniform sampling and strong curriculum baselines across a wide range of challenging reasoning benchmarks, demonstrating improved training stability and efficiency. Notably, it achieves relative gains of 28.6% on AIME2024 and 30.5% on ARC-1D over the strongest baseline and up to 80% speedup. These results suggest that ACTOR-CURATOR is a powerful and practical approach for scalable LLM post-training.
Zhengyao Gu, Jonathan Light, Raul Astudillo, Ziyu Ye, Langzhou He, Henry Peng Zou, Wei Cheng, Santiago Paternain, Philip S. Yu, Yisong Yue
arXiv:2602.20532 · cs.LG, cs.AI, cs.CL · submitted Feb 24, 2026
abstract · pdf · html · 37 pages, 8 figures, 1 table. Preprint under review. Equal contribution by first two authors
This paper looks at a problem that comes up in RL post-training of large models: the training data mixture (or curriculum) is often manually tuned and static, even though the policy keeps changing during training.
We propose Actor-Curator, a framework where a learned "curator" adaptively selects training problems while the actor policy is being optimized. The curator is trained to maximize a policy-improvement objective, effectively learning which data is most useful for improving the policy at each stage of training.
Conceptually it’s a co-adaptive system: - the actor learns the policy - the curator learns the training curriculum
Happy to answer questions or discuss!