about
Composing Linear Layers from Irreducibles (arxiv.org)
2 points by liamdgray on Aug 16, 2025 | hide | past | pdf | 1 comment on HN

In plain words: Any linear layer can be rebuilt by stacking simple rotations in pairs of directions, needing only about the square of the log of its width in parameters instead of the width squared. In language-model attention, these rotation-stacked layers matched the accuracy of strong compressed versions of those layers.

Abstract

Contemporary large models often exhibit behaviors suggesting the presence of low-level primitives that compose into modules with richer functionality, but these fundamental building blocks remain poorly understood. We investigate this compositional structure in linear layers by asking: can we identify/synthesize linear transformations from a minimal set of geometric primitives? Using Clifford algebra, we show that linear layers can be expressed as compositions of bivectors -- geometric objects encoding oriented planes -- and introduce a differentiable algorithm that decomposes them into products of rotors. This construction uses only O(log^2 d) parameters, versus O(d^2) required by dense matrices. Applied to the key, query, and value projections in LLM attention layers, our rotor-based layers match the performance of strong baselines such as block-Hadamard and low-rank approximations. Our findings provide an algebraic perspective on how these geometric primitives can compose into higher-level functions within deep models.

Travis Pence, Daisuke Yamada, Vikas Singh
arXiv:2507.11688 · cs.LG · submitted Jul 15, 2025 · updated Jun 9, 2026
abstract · pdf · html · 35 Pages, 11 Tables, 6 Figures, Appearing in NeurIPS 2025

add comment on HN

Abstract: "Contemporary large models often exhibit behaviors suggesting the presence of low-level primitives that compose into modules with richer functionality, but these fundamental building blocks remain poorly understood. We investigate this compositional structure in linear layers by asking: can we identify/synthesize linear transformations from a minimal set of geometric primitives? Using Clifford algebra, we show that linear layers can be expressed as compositions of bivectors -- geometric objects encoding oriented planes -- and introduce a differentiable algorithm that decomposes them into products of rotors. This construction uses only O(log^2 d) parameters, versus O(d^2) required by dense matrices. Applied to the key, query, and value projections in LLM attention layers, our rotor-based layers match the performance of strong baselines such as block-Hadamard and low-rank approximations. Our findings provide an algebraic perspective on how these geometric primitives can compose into higher-level functions within deep models."