about
Hierarchical Autoregressive Modeling for Memory-Efficient Language Generation (arxiv.org)
46 points by PaulHoule 271 days ago | hide | past | pdf | 3 comments on HN

In plain words: Instead of reading text one word at a time, this model squeezes tokens into coarse summaries and rebuilds fine details in parallel. It beats the usual one-word-at-a-time models on speed versus quality, with up to 1,000 times more output per unit of memory.

Abstract · PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation

Transformers operate as horizontal token-by-token scanners; at each generation step, attending to an ever-growing sequence of token-level states. This access pattern increases prefill latency and makes long-context decoding more memory-bound, as KV-cache reads and writes dominate inference time over arithmetic operations. We propose Parallel Hierarchical Operation for TOp-down Networks (PHOTON), a hierarchical autoregressive model that replaces horizontal scanning with vertical, multi-resolution context scanning. PHOTON maintains a hierarchy of latent streams: a bottom-up encoder compresses tokens into low-rate contextual states, while lightweight top-down decoders reconstruct fine-grained token representations in parallel. We further introduce recursive generation that updates only the coarsest latent stream and eliminates bottom-up re-encoding. Experimental results show that PHOTON is superior to competitive Transformer-based language models regarding the throughput-quality trade-off, providing advantages in long-context and multi-query tasks. In particular, this reduces decode-time KV-cache traffic, yielding up to $10^{3}\times$ higher throughput per unit memory.

Yuma Ichikawa, Naoya Takagi, Takumi Nakagawa, Yuzi Kanazawa, Akira Sakai
arXiv:2512.20687 · cs.LG, cs.AI, cs.CL, cs.DC · submitted Dec 22, 2025 · updated Jan 8, 2026
abstract · pdf · html · 17 pages, 10 figures

add comment on HN

At least the authors acknowledge it for what it is: a tiny model on a tiny corpus and worse than the comparable transformers in terms of accuracy. I like the experimentation with new designs and one doesnt always need to show near SOTA results. From a brief inspection, however, I think it will be hard for the work to become a high profile conference acceptance without significan additional work.
I would really like to see more testing with a deeper hierarchy and alpha and beta nonzero.
Skimming it I get this incredible sci-fi feeling of AI being the thing that solves P vs. NP (the diagrams are reminiscent of boolean/arithmetic circuits which have produced some results in the compcomp space)