about
DeepSeek Paper – DualPath: Breaking the Bandwidth Bottleneck in LLM Inference (arxiv.org)
2 points by dworks 220 days ago | hide | past | pdf | 1 comment on HN

In plain words: In multi-turn AI agents, conversation memory is pulled from storage into the engine that starts a reply, clogging its network link. This design routes it through the engine that writes replies first, then over a network, nearly doubling online serving speed without missing deadlines.

Abstract · DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference

The performance of multi-turn, agentic LLM inference is increasingly dominated by KV-Cache storage I/O rather than computation. In prevalent disaggregated architectures, loading the massive KV-Cache from external storage creates a fundamental imbalance: storage NICs on prefill engines become bandwidth-saturated, while those on decoding engines remain idle. This asymmetry severely constrains overall system throughput. We present DualPath, an inference system that breaks this bottleneck by introducing dual-path KV-Cache loading. Beyond the traditional storage-to-prefill path, DualPath enables a novel storage-to-decode path, in which the KV-Cache is loaded into decoding engines and then efficiently transferred to prefill engines via RDMA over the compute network. DualPath combines this optimized data path -- which inherently avoids network congestion and avoids interference with latency-critical model execution communications -- with a global scheduler that dynamically balances load across prefill and decode engines. Our evaluation on three models with production agentic workloads demonstrates that DualPath improves offline inference throughput by up to 1.87$\times$ on our in-house inference system. It can also improve online serving throughput by an average factor of 1.96$\times$ without violating SLO.

Yongtong Wu, Shaoyuan Chen, Yinmin Zhong, Rilin Huang, Yixuan Tan, Wentao Zhang, Liyue Zhang, Shangyan Zhou, Yuxuan Liu, Shunfeng Zhou, Mingxing Zhang, Xin Jin, et al.
arXiv:2602.21548 · cs.DC · submitted Feb 25, 2026 · updated Feb 26, 2026
abstract · pdf · html

add comment on HN
Also discussed: Jun 2026 (2 points, 0 comments) · Mar 2026 (2 points, 1 comment) · Feb 2026 (3 points, 0 comments)

Related: Why NVLink Is Nvidia’s Secret Sauce Driving a 10x Performance Boost in MoEs

https://www.hpcwire.com/2026/02/23/why-nvlink-is-nvidias-sec...