about
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models (arxiv.org)
33 points by dpstart01 on Apr 15, 2025 | hide | past | pdf | 3 comments on HN

In plain words: A math-reasoning model keeps memory flat as its thinking grows longer, trained to copy existing reasoners and improved with reinforcement learning. It matches same-size transformer reasoners and runs over 3 times faster, so voting across several answers beats them under a fixed time budget.

Abstract

Effective reasoning is crucial to solving complex mathematical problems. Recent large language models (LLMs) have boosted performance by scaling test-time computation through long chain-of-thought reasoning. However, transformer-based models are inherently limited in extending context length due to their quadratic computational complexity and linear memory requirements. In this paper, we introduce a novel hybrid linear RNN reasoning model, M1, built on the Mamba architecture, which allows memory-efficient inference. Our approach leverages a distillation process from existing reasoning models and is further enhanced through RL training. Experimental results on the AIME and MATH benchmarks show that M1 not only outperforms previous linear RNN models but also matches the performance of state-of-the-art Deepseek R1 distilled reasoning models at a similar scale. We also compare our generation speed with a highly performant general purpose inference engine, vLLM, and observe more than a 3x speedup compared to a same size transformer. With throughput speedup, we are able to achieve higher accuracy compared to DeepSeek R1 distilled transformer reasoning models under a fixed generation time budget using self-consistency voting. Overall, we introduce a hybrid Mamba reasoning model and provide a more effective approach to scaling test-time generation using self-consistency or long chain of thought reasoning.

Junxiong Wang, Wen-Ding Li, Daniele Paliotta, Daniel Ritter, Alexander M. Rush, Tri Dao
arXiv:2504.10449 · cs.LG · submitted Apr 14, 2025 · updated Sep 9, 2025
abstract · pdf · html · Code is available https://github.com/jxiw/M1

add comment on HN

Does anyone know if there were any attempts to test Mamba on really large scale? To me this model looks as the most promising successor to the transformer architecture. Does anyone know why it might not be the case or what are other alternatives?
Tencent's 'Hunyuan-T1'–The First Mamba-Powered Ultra-Large Model: https://news.ycombinator.com/item?id=43447254
Interesting direction for research but not a model you’d want to use today. The paper looks at a 3b model built on llama3.2-3b, modified for mamba, and they’re comparing to a distilled version of r1 with 1.5b params.