In plain words: By replaying real workloads from two companies, they tested 14 ways to decide which saved conversation beginnings to keep when memory fills. Fancy schemes barely beat keeping the most recent ones, since sessions return at steady intervals; small tweaks like dropping never-reused entries help most.
Abstract · When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood. We study production traces from two companies and evaluate 14 eviction algorithms across HBM-constrained and large memory-pool settings. Despite a large gap to Belady, sophisticated policies designed for traditional caches provide little benefit over LRU. The reason is structural: prefix reuse is dominated by the regular pacing of active sessions, making recency unusually predictive. Prefix caching nevertheless introduces new challenges, including heavy-tailed session footprints and highly variable miss costs as attention computation grows with sequence length. We introduce the compute-savings ratio and two offline oracles to quantify these effects. Our results show that effective prefix-cache management should retain recency as its foundation while selectively adding quick demotion for one-hit prefixes, compute-aware partial eviction for expensive misses, and capacity-dependent eviction granularity. We will release the traces and simulator to support future research.
Yiyu Liu, Minlan Yu, Juncheng Yang
arXiv:2609.28870 · cs.DC, cs.LG · submitted Sep 24, 2026 · updated Oct 1, 2026
abstract · pdf · html · 20 pages, 20 figures, 6 tables