about
Direct initialization of transformers using larger pretrained ones (arxiv.org)
48 points by PaulHoule on Dec 22, 2023 | hide | past | pdf | 14 comments on HN

In plain words: Weight subcloning shrinks a big pretrained transformer into a smaller one by keeping the most important neurons and dropping extra layers, giving it a smart start instead of random weights. It trained vision and language models about 4 times faster than starting from scratch.

Abstract · Weight subcloning: direct initialization of transformers using larger pretrained ones

Training large transformer models from scratch for a target task requires lots of data and is computationally demanding. The usual practice of transfer learning overcomes this challenge by initializing the model with weights of a pretrained model of the same size and specification to increase the convergence and training speed. However, what if no pretrained model of the required size is available? In this paper, we introduce a simple yet effective technique to transfer the knowledge of a pretrained model to smaller variants. Our approach called weight subcloning expedites the training of scaled-down transformers by initializing their weights from larger pretrained models. Weight subcloning involves an operation on the pretrained model to obtain the equivalent initialized scaled-down model. It consists of two key steps: first, we introduce neuron importance ranking to decrease the embedding dimension per layer in the pretrained model. Then, we remove blocks from the transformer model to match the number of layers in the scaled-down network. The result is a network ready to undergo training, which gains significant improvements in training speed compared to random initialization. For instance, we achieve 4x faster training for vision transformers in image classification and language models designed for next token prediction.

Mohammad Samragh, Mehrdad Farajtabar, Sachin Mehta, Raviteja Vemulapalli, Fartash Faghri, Devang Naik, Oncel Tuzel, Mohammad Rastegari
arXiv:2312.09299 · cs.LG, cs.CL, cs.CV · submitted Dec 14, 2023
abstract · pdf · html

add comment on HN

This might sound like nonsense to someone who actually understands neural networks, but I had a similar idea recently that perhaps we could use LLMs or diffusion models to generate the weights for a new network, e.g. if you think of a layer of a network as an image where each pixel represents the strength of a connection, then perhaps you could generate all the initial weights for a specialized network from a more general/powerful network. Curious if that’s been tried or if there are fundamental limitations that would prevent it.
I believe what you're looking for is Hypernetworks

https://arxiv.org/abs/1609.09106

Very relevant thanks!
Amazingly enough, that pretty much works, but you'll never hear about it because the name of it is impossible to search for: https://www.wpeebles.com/Gpt
Fun tiny fact I found a while ago - The reverse of this also works!

You can init deeper layers by copying the weights from previous layers. Not a big improvement in training time, but it does reduce the initial loss!

I would like to know if this approach is actually transferring knowledge, or if random initialization is just not the optimal starting place - for example, maybe having a different distribution of weights is better.

A good test would be to apply this technique, but then train on a totally different kind of data - for example take a text LLM, apply this technique, but then train on audio/music data and see if this technique reduces training time over random initialization.

Sounds like a good test.

We need something like this for "non-neural" things like transformers and normalization layers:

https://proceedings.mlr.press/v157/skorski21a/skorski21a.pdf

a better distribution of weights would be knowledge transfer right? by having something like a stronger prior distribution?

also pretty sure this would work because I know some text-image generators has the image network priority trained then frozen during e2e training

Not necessarily. It turns out to really matter how the weights are initialized (randomly, but from which distribution?) for ordinary feed-forward neural nets that only uses standard neurons -- different activation functions have different optimal initialization distributions.

https://machinelearningmastery.com/weight-initialization-for...

A big breakthrough was Kaiming He's work on initialization for ReLU:

https://arxiv.org/abs/1502.01852

(Kaiming is the personal name, He is the family name. He Kaiming is the native Chinese order. Lots of citations use "Kaiming" by mistake.)

This paper seems to me like it's trying to overstate its usefulness. I don't see the 4x reduction in training costs. And it also seems to require that the teacher and the student models have the same number of decoder layers, which for all common transformer models just isn't the case. It also doesn't work to train a transducer (a common real-time deployment architecture) based on a transformer teacher.

In short, I can't imagine ever having a training situation where this paper would be applicable.

Instead of pruning least contributing neurons wouldn't it be better to ie. nominate next lowest contributing neuron or most similar one to merge them into one or redistribute its impact onto others?
Pretty nice to get downscaled mobile models with will be huge in 2024 when the compilation platforms like onnx coreml webgpu get smooth enough
From Page 7:

> In the case of GPT-2 training, random initialization requires 64×10^9 tokens to reach a perplexity of 12, while weight subcloning accomplishes this in just 64×10^9 tokens, again demonstrating a 4× training speedup

... what?

Looks like typo, from chart it seems to be around 20x10^9 for weight subcloning (doesn't look like 16x10^9 <<which would be 4x>> either).