about
The End of Transformers (2025) (arxiv.org)
1 point by teleforce 246 days ago | hide | past | pdf | discuss on HN

In plain words: This survey compares sequence models—slimmer attention, recurrent networks, state-space models, and hybrids—against attention, whose compute and memory cost grows with the square of context length. The alternatives scale more cheaply on long texts, but each gives up something attention does well.

Abstract · The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures

Transformers have dominated sequence processing tasks for the past seven years -- most notably language modeling. However, the inherent quadratic complexity of their attention mechanism remains a significant bottleneck as context length increases. This paper surveys recent efforts to overcome this bottleneck, including advances in (sub-quadratic) attention variants, recurrent neural networks, state space models, and hybrid architectures. We critically analyze these approaches in terms of compute and memory complexity, benchmark results, and fundamental limitations to assess whether the dominance of pure-attention transformers may soon be challenged.

Alexander M. Fichtl, Jeremias Bohn, Josefin Kelber, Edoardo Mosca, Georg Groh
arXiv:2510.05364 · cs.CL · submitted Oct 6, 2025
abstract · pdf · html · 21 pages, 2 figures, 2 tables

add comment on HN