about
Mixture-of-Depths: Dynamically allocating compute in transformer language models (arxiv.org)
5 points by GaggiX on Apr 4, 2024 | hide | past | pdf | 2 comments on HN

In plain words: Instead of giving every word equal work at every layer, this model processes only the tokens each layer needs, keeping total work fixed. It matched the usual equal-work model's quality while using fewer calculations and running up to 50% faster when writing text.

Abstract · Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Transformer-based language models spread FLOPs uniformly across input sequences. In this work we demonstrate that transformers can instead learn to dynamically allocate FLOPs (or compute) to specific positions in a sequence, optimising the allocation along the sequence for different layers across the model depth. Our method enforces a total compute budget by capping the number of tokens ($k$) that can participate in the self-attention and MLP computations at a given layer. The tokens to be processed are determined by the network using a top-$k$ routing mechanism. Since $k$ is defined a priori, this simple procedure uses a static computation graph with known tensor sizes, unlike other conditional computation techniques. Nevertheless, since the identities of the $k$ tokens are fluid, this method can expend FLOPs non-uniformly across the time and model depth dimensions. Thus, compute expenditure is entirely predictable in sum total, but dynamic and context-sensitive at the token-level. Not only do models trained in this way learn to dynamically allocate compute, they do so efficiently. These models match baseline performance for equivalent FLOPS and wall-clock times to train, but require a fraction of the FLOPs per forward pass, and can be upwards of 50\% faster to step during post-training sampling.

David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, Adam Santoro
arXiv:2404.02258 · cs.LG, cs.CL · submitted Apr 2, 2024
abstract · pdf · html

add comment on HN
Also discussed: Apr 2024 (281 points, 83 comments) · Apr 2024 (2 points, 0 comments) · Apr 2024 (2 points, 1 comment) · Apr 2024 (4 points, 0 comments)

> MoD transformers demonstrate the value of routing among different types of computations. In this work the types were either the conventional transformer block, or a null computation (functionally equivalent to multiplying by zero). However, one can imagine extending this idea further by routing between even more types of computation. For example, perhaps some tokens are routed to "memory lookup" functions, and others are routed to "tool use" functions.

This is one of the most interesting parts of the paper for me. It seems that this could be a much more efficient approach for tool use and math calculation.

Pretty surprising that Google DeepMind is still publishing such papers unless they have something much further ahead already. I expect OpenAI and Anthropic to have their equivalents to this research, but it still lets others catch up more easily and makes LLMs closer to commodities.