about
Attention Once Is All You Need: Stateful Transformers (arxiv.org)
3 points by logotype 143 days ago | hide | past | pdf | 4 comments on HN

In plain words: Instead of re-reading history for every question, it keeps the model's memory alive and folds in new data as it arrives. On market data it ran up to 5.9 times faster than engines that rebuild context per query, with question time flat as history grew.

Abstract · Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers

Conventional transformer inference engines are request-driven, paying an O(n) prefill cost on every query. In streaming workloads, where data arrives continuously and queries probe an ever-growing context, this cost is prohibitive. We introduce a data-driven computational model centred on stateful sessions: a persistent KV cache advanced incrementally as new data arrives, so prefill is moved off the critical path and query latency becomes O(|q|), independent of accumulated context size. Building on this, Flash Queries reclaim idle GPU cycles between data arrivals to pre-evaluate registered questions and return cached answers before the user asks, a pattern that is structurally impossible in stateless engines because they discard intermediate state between requests. A multi-tenant continuous-batching scheduler with cell-budget admission and prefix-aware grouped prefill lets dozens of stateful sessions coexist on a single GPU while preserving full quadratic self-attention. On streaming market-data benchmarks the reference implementation achieves up to 5.9x speedup over conventional inference engines (vLLM, SGLang, TensorRT-LLM, llama.cpp), holding query latency constant as accumulated context grows.

Victor Norgren
arXiv:2605.13784 · cs.LG · submitted May 13, 2026
abstract · pdf · html

add comment on HN

Pin-drop sweepstakes
Happy to answer any questions you might have.
Tumbleweed rolls across the timeline
crickets