about
SpikingBrain Technical Spiking Brain-Inspired Large Models (arxiv.org)
3 points by knrz on Sep 17, 2025 | hide | past | pdf | 2 comments on HN

In plain words: A brain-inspired language model swaps the usual all-to-all word attention for linear attention and spiking nerve-like cells, so long texts need far less computing and memory. It matched standard models and began answering 100 times faster on 4-million-token inputs.

Abstract · SpikingBrain: Spiking Brain-inspired Large Models

Mainstream Transformer-based large language models face major efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly, limiting long-context processing. Building large models on non-NVIDIA platforms also poses challenges for stable and efficient training. To address this, we introduce SpikingBrain, a family of brain-inspired models designed for efficient long-context training and inference. SpikingBrain leverages the MetaX GPU cluster and focuses on three aspects: (1) Model Architecture: linear and hybrid-linear attention architectures with adaptive spiking neurons; (2) Algorithmic Optimizations: an efficient, conversion-based training pipeline and a dedicated spike coding framework; (3) System Engineering: customized training frameworks, operator libraries, and parallelism strategies tailored to MetaX hardware. Using these techniques, we develop two models: SpikingBrain-7B, a linear LLM, and SpikingBrain-76B, a hybrid-linear MoE LLM. These models demonstrate the feasibility of large-scale LLM development on non-NVIDIA platforms, and training remains stable for weeks on hundreds of MetaX GPUs with Model FLOPs Utilization at expected levels. SpikingBrain achieves performance comparable to open-source Transformer baselines while using only about 150B tokens for continual pre-training. Our models also significantly improve long-context efficiency and deliver inference with (partially) constant memory and event-driven spiking behavior. For example, SpikingBrain-7B attains over 100x speedup in Time to First Token for 4M-token sequences. Furthermore, the proposed spiking scheme achieves 69.15 percent sparsity, enabling low-power operation. Overall, this work demonstrates the potential of brain-inspired mechanisms to drive the next generation of efficient and scalable large model design.

Yuqi Pan, Yupeng Feng, Jinghao Zhuang, Siyu Ding, Han Xu, Zehao Liu, Bohan Sun, Yuhong Chou, Xuerui Qiu, Anlin Deng, Anjie Hu, Shurong Wang, et al.
arXiv:2509.05276 · cs.LG, cs.AI, cs.CL · submitted Sep 5, 2025 · updated May 8, 2026
abstract · pdf · html

add comment on HN
Also discussed: Sep 2025 (1 point, 0 comments)

"Researchers get spiking neural behavior out of a pair of transistors" (2025) https://news.ycombinator.com/item?id=43506198

What are the ways to get spiking behavior out of integrated nanophotonics?

Saturable Absorption (excitable semiconductor lasers, graphene laser cavity,), NDR Negative Differential Resistance (RTD Resonant Tunneling Diodes,), PCM: Phase-change materials (DVD-RW,),

Metamaterials and metasurfaces are probably useful for extreme nonlinear spiking neuromorphic computing with integrated nanophotonics.

What about Optical rogue waves, Supercontinuum generation (color, wave division multiplexing, ); and/or Superradiance (as nonlinear optical effects for a neuromorphic computation platform)? ... https://news.ycombinator.com/item?id=41684444

Superradiance: https://en.wikipedia.org/wiki/Superradiance :

> Superradiance has since been demonstrated in a wide variety of physical and chemical systems, such as quantum dot arrays [4] and J-aggregates. [5] This effect has been used to produce a superradiant laser.

Superradiance in semiconductor optics -> Coherent effects in semiconductor optics > Superradiance of excitons: https://en.wikipedia.org/wiki/Coherent_effects_in_semiconduc...