about
Surgical Repair of Collapsed Attention Heads in ALiBi Transformers (arxiv.org)
3 points by palmerschallon 207 days ago | hide | past | pdf | 2 comments on HN

In plain words: A built-in distance penalty makes many attention heads in these models stare almost only at the first token; the fix resets those heads and freezes the rest. It restored 98.7% of heads (242 to 379 of 384) in two passes on one consumer GPU.

Abstract

We identify a systematic attention collapse pathology in the BLOOM family of transformer language models, where ALiBi positional encoding causes 31-44% of attention heads to attend almost entirely to the beginning-of-sequence token. The collapse follows a predictable pattern across four model scales (560M to 7.1B parameters), concentrating in head indices where ALiBi's slope schedule imposes the steepest distance penalties. We introduce surgical reinitialization: targeted Q/K/V reinitialization with zeroed output projections and gradient-masked freezing of all non-surgical parameters. Applied to BLOOM-1b7 on a single consumer GPU, the technique recovers 98.7% operational head capacity (242 to 379 of 384 heads) in two passes. A controlled comparison with C4 training data confirms that reinitialization -- not corpus content -- drives recovery, and reveals two distinct post-surgical phenomena: early global functional redistribution that improves the model, and late local degradation that accumulates under noisy training signal. An extended experiment reinitializing mostly-healthy heads alongside collapsed ones produces a model that transiently outperforms stock BLOOM-1b7 by 25% on training perplexity (12.70 vs. 16.99), suggesting that pretrained attention configurations are suboptimal local minima. Code, checkpoints, and diagnostic tools are released as open-source software.

Palmer Schallon
arXiv:2603.09616 · cs.CL · submitted Mar 10, 2026
abstract · pdf · html · 15 pages, 7 figures, 2 supplementary figures. Code: https://github.com/Palmerschallon/bloom-head-surgery Checkpoints: https://huggingface.co/TheNexus42/bloom-1b7-head-surgery

add comment on HN

Interesting that lack of engineering discipline in LLM design has industry vernacular veering into medically pathologized forensics and remediation.

From a C.A.R. Hoare perspective, this looks like a horrible development: Human bias in training of LLMs produces effects that convince credulous users that automata are alive, credulous users think that as we don't understand why living organisms do what they do, then as an LLM seems like a living organism there's no expectation of understanding why LLMs do what they do...

We put our own ghost in a machine, confuse the machine with our ghost, then rely on dark arts to cope with our lack of understanding of the machine.

So expect LLMs to be further mystified, treated as specimens for study and symptoms to cure, then categorized behaviorally by medically appropriated syndromes.

We can already see the burgeoning new high priest career-track Doctor of AI, with all the attendant quackery, patent cures, leaching and horse-whispering.

Found that ALiBi positional encoding causes 31-44% of attention heads in BLOOM-family models to collapse — attending almost entirely to token 0 rather than meaningful context. The paper identifies the pathology and a targeted repair. Happy to answer questions.