about
Log-Depth Recurrent Language Modeling (arxiv.org)
1 point by E-Reverance 9 days ago | hide | past | pdf | discuss on HN

In plain words: A language model folds text into a balanced tree of combined chunks, so predicting the next word takes slowly growing steps and linear work, unlike Transformers' fixed depth and quadratic cost. It handled far longer texts than trained on and nearly matched a Transformer.

Abstract

Language modeling using Transformers has become commonplace despite their fixed computational depth and quadratic runtime with respect to input tokens. Recurrent models on the other hand offer linear depth but no parallel execution. In this work, we extend balanced-tree recursive operators from sequence encoding to autoregressive prediction, enabling all prefix representations to be computed with logarithmic depth and linear runtime. Our experiments provide an initial characterization of this model class, demonstrating robust length extrapolation and performance approaching that of ALiBi-based Transformers, highlighting its potential as an alternative architecture for language modeling.

Yiqin Wang, Nuri Cingillioglu, Charles Pert
arXiv:2609.28212 · cs.LG, cs.CL · submitted Sep 23, 2026
abstract · pdf · html · 5 pages, 3 figures

add comment on HN