about
Sliding Window Recurrences for Sequence Models (arxiv.org)
1 point by nsoonhui 271 days ago | hide | past | pdf | discuss on HN

In plain words: A new layer chops a running memory of past words into jagged, hardware-friendly chunks so the chip's processing groups rarely talk to each other. In 1-billion-parameter language models it ran 10-40% faster than optimized Transformers while guessing the next word just as well.

Abstract

Multi-hybrid architectures are poised to take over language modeling due to better quality and performance. We introduce a hierarchical decomposition framework for linear recurrences that allows us to develop algorithms aligned with GPU memory hierarchies, yielding Sliding Window Recurrences. We focus specifically on truncating recurrences to hardware-aligned windows which are naturally jagged, limiting costly inter-warp communication. Using SWR, we develop Phalanx layers that serve as drop-in replacements for windowed attention or linear recurrences. In 1B parameter multi-hybrid models, Phalanx achieves over 10-40% speedup across 4K to 32K context length over optimized Transformers while matching perplexity.

Dragos Secrieru, Garyk Brixi, Yoshua Bengio, Taiji Suzuki, Michael Poli, Stefano Massaroli
arXiv:2512.13921 · cs.LG · submitted Dec 15, 2025
abstract · pdf

add comment on HN