about
Brainformers: Trading Simplicity for Efficiency (arxiv.org)
81 points by PaulHoule on Jun 3, 2023 | hide | past | pdf | 3 comments on HN

In plain words: Instead of repeating the same attention-then-feedforward pattern, this design mixes different layer types in varied orders inside each block. It reached good results twice as fast as a comparable model and scored higher on language-understanding tests.

Abstract

Transformers are central to recent successes in natural language processing and computer vision. Transformers have a mostly uniform backbone where layers alternate between feed-forward and self-attention in order to build a deep network. Here we investigate this design choice and find that more complex blocks that have different permutations of layer primitives can be more efficient. Using this insight, we develop a complex block, named Brainformer, that consists of a diverse sets of layers such as sparsely gated feed-forward layers, dense feed-forward layers, attention layers, and various forms of layer normalization and activation functions. Brainformer consistently outperforms the state-of-the-art dense and sparse Transformers, in terms of both quality and efficiency. A Brainformer model with 8 billion activated parameters per token demonstrates 2x faster training convergence and 5x faster step time compared to its GLaM counterpart. In downstream task evaluation, Brainformer also demonstrates a 3% higher SuperGLUE score with fine-tuning compared to GLaM with a similar number of activated parameters. Finally, Brainformer largely outperforms a Primer dense model derived with NAS with similar computation per token on fewshot evaluations.

Yanqi Zhou, Nan Du, Yanping Huang, Daiyi Peng, Chang Lan, Da Huang, Siamak Shakeri, David So, Andrew Dai, Yifeng Lu, Zhifeng Chen, Quoc Le, et al.
arXiv:2306.00008 · cs.LG, cs.CL · submitted May 29, 2023 · updated Apr 25, 2024
abstract · pdf · html

add comment on HN

The Brainformer building block is designed using neural architecture search.

"Brainformer consistently outperforms the state-of-the-art dense and sparse Transformers, in terms of both quality and efficiency. A Brainformer model with 8 billion activated parameters per token demonstrates 2× faster training convergence and 5× faster step time compared to its GLaM counterpart."

Note that the 8B model mentioned above has 158B total parameters. The authors compare the training time to other sparsely activated models, but seemingly not to dense models.

I think parsimony is great for human readability but also creates a bias that excludes within reach solutions that overall have meaningfully better properties https://en.wikipedia.org/wiki/Evolved_antenna
Architecture search always works for improving efficiency, but it never beats a real algorithmic improvement (and there were lots of interesting ones being proposed for increasing context window / improving inference time in the last few months).