In plain words: Turns speech tokens into sound in real time by letting each chunk of audio listen only to its neighboring chunks instead of the whole recording. It matched the quality of offline versions, beat other real-time ones, and started playing after 180 milliseconds.
Abstract · StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
Recent advancements in discrete token-based speech generation have highlighted the importance of token-to-waveform generation for audio quality, particularly in real-time interactions. Traditional frameworks integrating semantic tokens with flow matching (FM) struggle with streaming capabilities due to their reliance on a global receptive field. Additionally, directly implementing token-by-token streaming speech generation often results in degraded audio quality. To address these challenges, we propose StreamFlow, a novel neural architecture that facilitates streaming flow matching with diffusion transformers (DiT). To mitigate the long-sequence extrapolation issues arising from lengthy historical dependencies, we design a local block-wise receptive field strategy. Specifically, the sequence is first segmented into blocks, and we introduce block-wise attention masks that enable the current block to receive information from the previous or subsequent block. These attention masks are combined hierarchically across different DiT-blocks to regulate the receptive field of DiTs. Both subjective and objective experimental results demonstrate that our approach achieves performance comparable to non-streaming methods while surpassing other streaming methods in terms of speech quality, all the while effectively managing inference time during long-sequence generation. Furthermore, our method achieves a notable first-packet latency of only 180 ms.\footnote{Speech samples: https://dukguo.github.io/StreamFlow/}
Dake Guo, Jixun Yao, Linhan Ma, He Wang, Lei Xie
arXiv:2506.23986 · cs.SD, eess.AS · submitted Jun 30, 2025 · updated Jul 1, 2025
abstract · pdf · html
Why this matters: Current diffusion speech models need to see the entire audio sequence, making them too slow and memory-heavy for assistants, agents, or anything that needs instant voice responses. Causal masks sound robotic; chunking adds weird seams. Streaming TTS has been stuck with a quality–latency tradeoff.
The idea: StreamFlow restricts attention using sliding windows over blocks:
Each block can see W_b past blocks and W_f future blocks
Compute becomes roughly O(B × W × N) instead of full O(N²)
Prosody stays smooth, latency stays constant, and boundaries disappear with small overlaps + cross-fades
How it works: The system is still a Diffusion Transformer, but trained in two phases:
Full-attention pretraining for global quality
Block-wise fine-tuning to adapt to streaming constraints
Generates mel-spectrograms; BigVGAN vocoder runs in parallel.
Performance:
~180ms first-packet latency (80ms model, 60ms vocoder, 40ms overhead)
No latency growth with longer speech
MOS tests show near-indistinguishable quality vs non-streaming diffusion
Speaker similarity within ~2%, prosody continuity preserved
Key ablation takeaways:
Past context helps until ~3 blocks; more adds little
Even a tiny future window greatly boosts naturalness
Best results: 0.4–0.6s block size, ~10–20% overlap
Comparison:
Autoregressive TTS → streaming but meh quality
GAN TTS → fast but inconsistent
Causal diffusion → real-time but degraded
StreamFlow → streaming + near-SOTA quality
Bigger picture: Smart attention shaping lets diffusion models work in real time without throwing away global quality. The same technique could apply to streaming music generation, translation, or interactive media.