In plain words: A chatbot saves the math from reading a prompt so follow-ups come fast; today each cloud keeps its copy. This vision shares those saved chunks across clouds, deciding where to store or recompute them by network speed and price to cut delay and cost.
Abstract · An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age
LLM inference has become a global-scale, heterogeneous workload spanning agents, retrieval, tool-use, code execution and multi-modal reasoning. These workloads naturally enable context reuse from overlapping inputs, creating a major opportunity to store and reuse the contexts' KV Caches instead of recomputing them. However, model-side advances that shrink the KV Cache and system-side advances that reduce compute, storage, and transfer costs are evolve independently within legacy cloud boundaries. We argue that future inference infrastructure should allow decoupling of compute and KV Cache storage across cloud and datacenters. The network becomes an active distribution channel; bandwidth, latency and pricing directly determines how the KV Cache should be managed. We propose a vision for an Internet for the KV Cache, with KV Cache management working as a content-distribution system. In this view, KV Cache storage and recompute decisions are driven by model, infrastructure, and application metrics, to enable adaptive, content-driven decisions for minimizing latency and cost.
Siddhant Ray, Nick Feamster, Junchen Jiang
arXiv:2608.01526 · cs.NI · submitted Aug 2, 2026
abstract · pdf · html