about
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning (arxiv.org)
1 point by lexandstuff 145 days ago | hide | past | pdf | discuss on HN

In plain words: Skill1 trains one agent to search a library of saved strategies, pick one, use it to finish a task, and save a new one, all rewarded only by task success. It beat earlier skill-reusing and trial-and-error agents in two simulated home and web tasks.

Abstract

A persistent skill library allows language model agents to reuse successful strategies across tasks. Maintaining such a library requires three coupled capabilities. The agent selects a relevant skill, utilizes it during execution, and distills new skills from experience. Existing methods optimize these capabilities in isolation or with separate reward sources, resulting in partial and conflicting evolution. We propose Skill1, a framework that trains a single policy to co-evolve skill selection, utilization, and distillation toward a shared task-outcome objective. The policy generates a query to search the skill library, re-ranks candidates to select one, solves the task conditioned on it, and distills a new skill from the trajectory. All learning derives from a single task-outcome signal. Its low-frequency trend credits selection and its high-frequency variation credits distillation. Experiments on ALFWorld and WebShop show that Skill1 outperforms prior skill-based and reinforcement learning baselines. Training dynamics confirm the co-evolution of the three capabilities, and ablations show that removing any credit signal degrades the evolution.

Yaorui Shi, Yuxin Chen, Zhengxi Lu, Yuchun Miao, Shugui Liu, Qi GU, Xunliang Cai, Xiang Wang, An Zhang
arXiv:2605.06130 · cs.AI · submitted May 7, 2026 · updated May 12, 2026
abstract · pdf · html

add comment on HN