In plain words: A controller shrinks how much new prompt text is processed at once whenever it overlaps active decoding, using the gap between scheduling rounds as feedback to size the next chunk. It cut worst-case token gaps by about 28% on a small model, but failed on larger and multi-GPU setups because that timing gap misrepresents actual GPU work.
Abstract · Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency objective. We introduce Decode-Latency Feedback Prefill (DLFP), a model-free controller that changes only prefill work that overlaps active decodes. After a guarded scheduling cycle, DLFP uses the observed interval as proportional feedback to resize the next prefill chunk; isolated prefills remain unrestricted. We implement DLFP in vLLM and evaluate it with open-loop Poisson arrivals, exact token accounting, raw request traces, and NVIDIA telemetry. On Qwen3-0.6B in BF16 on one A100 80 GB GPU, three paired 100-request trials reduce P99 inter-token latency by 24.8%, 30.1%, and 28.2% (mean 27.7%, paired 95% confidence interval 21.0% to 34.3%) with exact output agreement, no failures, and unchanged SLO compliance. The benefit is not free: mean P99 time to first token increases 34.8% while remaining inside the declared SLO. Crucially, the mechanism does not generalize to Qwen3-8B, Qwen3-32B, or a two-GPU tensor-parallel configuration. We trace the failure to an asynchronous scheduler-call interval that is only a proxy for completed GPU iteration time. This negative result defines the boundary of the contribution and motivates a completion-timed controller for concurrent CPU and on-device inference. We do not claim mobile-device performance; the present work is a reproducible proof-of-concept and generalization study.
Gaurav Agarwal, Ashish Garg, Isha Singhal
arXiv:2609.38386 · cs.AI · submitted Sep 29, 2026
abstract · pdf · html · 5 pages, 1 figure. Includes negative generalization results for Qwen3-8B, Qwen3-32B, and two-GPU tensor parallelism
Happy to answer questions about the implementation, experiments, or results.