about
Prefill-as-a-Service:KVCache of Next-Generation Models Could Go Cross-Datacenter (arxiv.org)
44 points by matt_d 168 days ago | hide | past | pdf | 1 comment on HN

In plain words: Newer models shrink the cache passed from prompt reading to answer writing, so long requests can be handled in one datacenter and shipped over Ethernet to another. In a test with a huge model, it handled 54% more traffic than the usual single-datacenter setup.

Abstract · Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter

Prefill-decode (PD) disaggregation has become the standard architecture for large-scale LLM serving, but in practice its deployment boundary is still determined by KVCache transfer. In conventional dense-attention models, prefill generates huge KVCache traffics that keep prefill and decode tightly coupled within a single high-bandwidth network domain, limiting heterogeneous deployment and resource elasticity. Recent hybrid-attention architectures substantially reduce KVCache size, making cross-cluster KVCache transport increasingly plausible. However, smaller KVCache alone does not make heterogeneous cross-datacenter PD serving practical: real workloads remain bursty, request lengths are highly skewed, prefix caches are unevenly distributed, and inter-cluster bandwidth fluctuates. A naive design that fully externalizes prefill can therefore still suffer from congestion, unstable queueing, and poor utilization. We present Prefill-as-a-Service (PrfaaS), a cross-datacenter serving architecture that selectively offloads long-context prefill to standalone, compute-dense prefill clusters and transfers the resulting KVCache over commodity Ethernet to local PD clusters for decode. Rather than treating reduced KVCache as sufficient, PrfaaS combines model-side KV efficiency with system-side selective offloading, bandwidth-aware scheduling, and cache-aware request placement. This design removes the requirement that heterogeneous accelerators share the same low-latency RDMA fabric, enabling independent scaling of prefill and decode capacity across loosely coupled clusters. In a case study using an internal 1T-parameter hybrid model, a PrfaaS-augmented heterogeneous deployment achieves 54% higher serving throughput and 64% lower P90 TTFT than a homogeneous PD baseline, with approximately 15% throughput gain at equal cost, while consuming only modest cross-datacenter bandwidth.

Ruoyu Qin, Weiran He, Yaoyu Wang, Zheming Li, Xinran Xu, Yongwei Wu, Weimin Zheng, Mingxing Zhang
arXiv:2604.15039 · cs.DC · submitted Apr 16, 2026 · updated Apr 22, 2026
abstract · pdf · html · 16 pages, 5 figures, 6 tables

add comment on HN

Maybe I'm missing something in this paper, but this seems to me to be just pretty "standard" caching stuff, albeit:

a) very time sensitive b) huge files c) scoped per user

Sort of reminds me of video streaming on CDNs for live video (but per user)?

I still think the big win is going to come based on time of use/live capacity. In a pure economics sense you want to charge a lot for inference when it's oversubscribed and far less when it's off peak (see electricity markets).

We have seen this with anthropics peak times, but it's very blunt currently. We also saw this with batch processing back in the day, but that breaks down because agents are 'chatty' and need to send new responses ASAP. You can't wait ages for each response - it would take weeks to do a simple agentic task if you had to wait hours between turn.

So I think what we'll see is async agents queued up, that you can then decide when to run them - either 'immediately' for time sensitive stuff (for more $$$) or 'best effort' where they can be scheduled to run whenever the provider wants to (3am say). If you also have diagnostics that usually agent task xyz takes y tokens total you can do far more efficient scheduling of these. This also reduces the amount of KVcache gymnastics significantly, as you can dedicate that agent task to a certain rack and schedule it all efficiently.

tl;dr I think the issues with inference efficiency need to be solved at a higher abstraction level of per agent "task" not purely on a per chat message basis. If you can schedule a load of agentic use cases off peak you don't need to preempt them because there is spare capacity by nature.