about
MoonshotAI unveils Kimi's large-scale LLM serving architecture (arxiv.org)
18 points by slothfulhamster on Jul 2, 2024 | hide | past | pdf | 1 comment on HN

In plain words: Mooncake splits a chat service so one group of machines reads your prompt and another writes the reply, reusing the model's saved state of earlier text kept in spare memory and drives. It lets Kimi handle 75% more requests than the usual setup while meeting speed targets.

Abstract · Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume all requests will be processed, Mooncake faces challenges due to highly overloaded scenarios. To mitigate these, we developed a prediction-based early rejection policy. Experiments show that Mooncake excels in long-context scenarios. Compared to the baseline method, Mooncake can achieve up to a 525% increase in throughput in certain simulated scenarios while adhering to SLOs. Under real workloads, Mooncake's innovative architecture enables Kimi to handle 75% more requests.

Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu
arXiv:2407.00079 · cs.DC, cs.AI, cs.AR · submitted Jun 24, 2024 · updated Sep 3, 2025
abstract · pdf · html · 23 pages, 13 figures

add comment on HN

I have been wondering the reason why online generative AI can serving so many requests. This really gives me an explanation.