In plain words: A new web layer would cache and serve short, self-contained passages instead of whole pages, letting many AI agents share search and processing work while sites keep control. Early tests found most page content goes unused, passages get reused, and answers improve per token.
Abstract · Semantics Delivery Network: Rethinking Web Retrieval Infrastructure for LLM Agents
Large language models (LLMs) increasingly rely on external sources when answering questions that require proprietary information or up-to-date live web content, through both traditional single-shot retrieval-augmented generation (RAG) and multi-turn agentic RAG. Yet today's web infrastructure is still built for human clients. Given a query, current search services return a list of URLs and snippets ranked for generic relevance; content delivery networks (CDNs) cache URL-addressed objects (texts, images, videos, etc.) without knowing which passage an agent needs. LLMs, in contrast, consume short, semantically coherent passages, hereafter "chunks", selected for downstream task utility rather than similarity alone, and may retrieve statefully across reasoning turns. Uncoordinated agents also repeat search, data acquisition, and semantic processing, duplicating work that could be shared. We argue that semantic chunk retrieval should become a first-class network-delivery abstraction. We propose Semantics Delivery Network (SemDN): an origin-authorized, hierarchical edge substrate that indexes, searches, and smart-caches web content at chunk granularity. SemDN serves agents on behalf of participating websites, amortizes data acquisition and processing across agents, and supports tenant-specific retrieval policies. Because, unlike URL caching, semantic retrieval provides no explicit miss signal, SemDN must estimate when its enrolled corpus may be incomplete or stale and trigger scoped discovery or refresh. It raises open questions about shareable retrieval state, hierarchical caching, coverage risk, and deployment. Our preliminary probes reveal a large gap between page content processed and chunks consumed, substantial task-local reuse, and higher answer quality per context token from chunk delivery.
Peichun Hua, Yunming Xiao
arXiv:2609.22486 · cs.NI, cs.IR, cs.LG · submitted Sep 18, 2026
abstract · pdf · html · 12 pages, 3 figures
This paper proposes a different approach:
1. Use a query retrieval system to push only the relevant paragraphs to the AI. 2. Consolidate and deduplicate the information from those sections.
This approach cuts costs and runs much faster.
This method solves the efficiency problem, but accuracy cannot be guaranteed. It relies entirely on the capability of the query system mentioned in the first point here.
Another way to improve data accuracy is to extract and transform the semantic layer before data entry. This does make writing data slower, but semantic data is different from transactional business data.
Semantic data doesn't necessarily need to be consumed by operational systems immediately. From a database perspective, it doesn't require OLTP characteristics. Therefore, a better approach might be to let this data undergo a period of processing and cleaning to become meaningful semantic information before making it available for subsequent retrieval.