In plain words: Instead of rewarding correct answers using gold solutions, the model scores its own confidence in each answer and uses that as its only reward. It matched standard reward-based training on math and did better on code, needing no labels or test cases.
Abstract · Learning to Reason without External Rewards
Training large language models (LLMs) for complex reasoning via Reinforcement Learning with Verifiable Rewards (RLVR) is effective but limited by reliance on costly, domain-specific supervision. We explore Reinforcement Learning from Internal Feedback (RLIF), a framework that enables LLMs to learn from intrinsic signals without external rewards or labeled data. We propose Intuitor, an RLIF method that uses a model's own confidence-termed self-certainty-as its sole reward signal. Intuitor replaces external rewards in Group Relative Policy Optimization (GRPO) with self-certainty scores, enabling fully unsupervised learning. Experiments demonstrate that Intuitor matches GRPO's performance on mathematical benchmarks while achieving better generalization to out-of-domain tasks like code generation, without requiring gold solutions or test cases. Our findings show that intrinsic model signals can drive effective learning across domains, offering a scalable alternative to RLVR for autonomous AI systems where verifiable rewards are unavailable. Code is available at https://github.com/sunblaze-ucb/Intuitor
Xuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine, Dawn Song
arXiv:2505.19590 · cs.LG, cs.CL · submitted May 26, 2025 · updated May 16, 2026
abstract · pdf · html · ICLR 2026