about
Tool-schema compression enables agentic RAG under constrained context budgets (arxiv.org)
2 points by Sakizli 130 days ago | hide | past | pdf | 1 comment on HN

In plain words: Tool descriptions eat the space a chatbot needs for search results, so this study shrinks them by about half. With a tiny 8K window, full descriptions overflowed and accuracy fell to 2.6%; the shrunken ones restored it, gaining 20.5 points.

Abstract · Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets

Agentic RAG systems that equip language models with dozens to hundreds of tool definitions face a critical resource conflict: tool schemas consume the same context window needed for retrieval-augmented generation. We present the first systematic study of this tool-context trade-off, evaluating 14 models spanning 1.5B-32B local models plus one frontier API model across 6,566 controlled API calls at three context budgets (8K, 16K, 32K) with 28 tool definitions. Applying TSCG conservative-profile compression (44-50% schema token savings), we observe a binary enablement effect: at 8K tokens, JSON-schema tool definitions overflow the context window entirely, yielding near-zero EM (2.6% average), while compressed schemas restore RAG functionality with +20.5 pp average exact-match lift across all eight models (+24.7 pp among the six exhibiting full enablement). At 32K -- where both formats fit -- four of five tested models show delta <= 1 pp, confirming the effect is purely budget-driven. External validation on HotpotQA (50 multi-hop questions) shows +48 pp EM under the same overflow scenario. Frontier scaling tests demonstrate that JSON schemas overflow at ~494 tools while compressed schemas remain operational beyond 800 tools. Our results establish tool-schema compression as a necessary infrastructure layer for agentic RAG in constrained-context deployments. All code, data, and checkpoints are publicly available.

Furkan Sakizli
arXiv:2605.26165 · cs.SE, cs.AI, cs.CL · submitted May 24, 2026
abstract · pdf · html · 12 pages (8 main + 4 appendix), 7 tables, 2 figures. Code and data: https://github.com/SKZL-AI/tscg

add comment on HN

How to best manage a large and growing collection of tools is something I've been fighting with a lot recently.

I am currently finding that progressive discovery via something like a Tools table in SQL might be the best option. The "compression" ratio achieved here can be extreme. You could have a million tools and as long as only a handful match the query each time, it still wouldn't overwhelm the context window. You can also add biz rules to this process. For example, you can force the agent to further constrain its tool search if too many rows are returned.