about
Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use (arxiv.org)
2 points by Anon84 218 days ago | hide | past | pdf | discuss on HN

In plain words: A training setup teaches a model to rewrite tool descriptions so agents pick the right tool, learning from worked examples first and then without them so it handles new APIs without extra runs. With 150+ tools, it cut the usual accuracy drop by 29%.

Abstract

While most efforts to improve LLM-based tool-using agents focus on the agent itself - through larger models, better prompting, or fine-tuning - agent performance increasingly plateaus due to the quality of the tool interfaces these agents consume. Tool descriptions are often written for human developers and tolerate ambiguity that agents cannot resolve, particularly as the number of candidate tools grows. Existing approaches to improving tool interfaces (1) require re-running a multi-stage per-tool pipeline - synthesizing queries, executing an agent to collect trajectories, annotating trajectories, and prompting a strong LLM multiple times - for every API that enters the catalog, and (2) typically optimize each tool independently, limiting scalability and generalization to unseen tools. We propose Trace-Free+, a curriculum learning framework that progressively transfers supervision from trace-rich settings to trace-free deployment, encouraging the model to internalize reusable patterns of what makes a tool description effective. To support this approach, we construct a large-scale dataset of high-quality tool interfaces derived from real-world APIs through a principled data synthesis workflow. Experiments on widely adopted benchmarks show that Trace-Free+ improves robustness as tool catalogs scale to 150+ candidates - in scaling experiments, reducing accuracy degradation by 29.23% and improving average query-level success by 60.89% on StableToolBench - generalizes across domains without retraining, and provides complementary gains on top of agent fine-tuning.

Ruocheng Guo, Kaiwen Dong, Xiang Gao, Kamalika Das
arXiv:2602.20426 · cs.AI · submitted Feb 23, 2026 · updated Apr 29, 2026
abstract · pdf · html · Preprint

add comment on HN