about
Expanding Transformer size without losing function or starting from scratch (arxiv.org)
52 points by famouswaffles on Aug 18, 2023 | hide | past | pdf | 26 comments on HN

In plain words: Instead of retraining a bigger network from scratch, these transformations grow a transformer's size while keeping its output exactly the same, so capacity can be added during training. Each expansion is proven to preserve the function perfectly when the new weights are set up in a minimal way.

Abstract · Composable Function-preserving Expansions for Transformer Architectures

Training state-of-the-art neural networks requires a high cost in terms of compute and time. Model scale is recognized to be a critical factor to achieve and improve the state-of-the-art. Increasing the scale of a neural network normally requires restarting from scratch by randomly initializing all the parameters of the model, as this implies a change of architecture's parameters that does not allow for a straightforward transfer of knowledge from smaller size models. In this work, we propose six composable transformations to incrementally increase the size of transformer-based neural networks while preserving functionality, allowing to expand the capacity of the model as needed. We provide proof of exact function preservation under minimal initialization constraints for each transformation. The proposed methods may enable efficient training pipelines for larger and more powerful models by progressively expanding the architecture throughout training.

Andrea Gesmundo, Kaitlin Maile
arXiv:2308.06103 · cs.LG · submitted Aug 11, 2023
abstract · pdf · html

add comment on HN

I’m surprised we weren’t doing this already. I’d like to see what happens if you train a small language model on preschool level reading material, and ramp up both the model size and training data complexity as you go. My hope would be you’d need less data to train a model in this fashion than you would with our current approach of throwing the entire internet at the model.
You might be interested in TinyStories:

https://arxiv.org/abs/2305.07759

> In this work, we introduce TinyStories, a synthetic dataset of short stories that only contain words that a typical 3 to 4-year-olds usually understand, generated by GPT-3.5 and GPT-4. We show that TinyStories can be used to train and evaluate LMs that are much smaller than the state-of-the-art models (below 10 million total parameters), or have much simpler architectures (with only one transformer block), yet still produce fluent and consistent stories with several paragraphs that are diverse and have almost perfect grammar, and demonstrate reasoning capabilities.

> I’m surprised we weren’t doing this already.

Because they don't work very well in real life. Article doesn't have any experimental proof it works, saves training time, or size, or anything..

And it doesn't mention another promising approach: mix of experts. There are many ways of implementing it. I mean the general idea of specialized fragments of the NN which are selectively called. They don't have to be trained all at once.

It’s weird a DeepMind author is on this as they experimentally studies these methods in Gopher and found it was useless.

Even the layer duplication trick they were using doesn’t work. I tried a refinement of that where you fine tuned the activations of the added layers to match the original activations which helped a bit relative to what DeepMind found but it still wasn’t material enough to be worth bothering with

I tried adding layer. The conclusion is, it's useful if you cannot train the whole model at once. Then you can split it in 2 and train the first half, with embedding/deembedding layers. After it stops improving add the second half of the layers, and train only it, freezing everything else. This worked in my case, when model params number < 10% of the number of training tokens. Splitting model in 3 is significantly less effective. Don't know why, but the third part adds almost nothing in accuracy and a lot in training time.

One thing where additive training works is fine tuning, like LoRA. But that's different. It's like poor man's way. For those who cannot train the whole model. Most likely uptrainig it whole is more efficient, just my guess.

The general term for this idea is curriculum learning.
I actually authored a patent for something related to this in 2020: https://scholar.google.com/citations?view_op=view_citation&h...

Not specifically for Transformers but for all kinds of neural networks.

It's about training long-lived neural networks where the training spans longer time scales than the neural network architecture remains valid. For example in self-driving cars, it's important to be able to train a model with collected live data over the years so that we can upgrade the model architecture as we go without throwing out the training already done.

The patent describes operations where the model architecture can be mutated step by step from old architecture to a new one without forgetting too much of the training which was applied to the old model. This includes but isn't limited to: Increasing/decreasing the width of neuronal layers, adding/removing layers, adding/removing inputs, adding/removing outputs, changing activation functions, changing loss functions. It also describes planning the steps of how to migrate from the old architecture to the new one.

Cool patent, looks like it covers everything.
Do I understand correctly that they are showing, "if we increase the size of these tensors along these directions, by padding them with zeros, it does not change the inference behavior of the network."?

If so, I'm.. kind of surprised that this wasn't found earlier?

Ah, well, they also show that they can also insert a new layer in the middle, but, because the layers output changes to be added on to the vector that was received, it does not seem terribly surprising to me that they can initialize a layer such that it produces 0 as the vector it should add onto it.

Ah, wait, no, they also show that other parts of tensors that are added when increasing the size of some parameters, can be initialized arbitrarily, provided that the appropriate other parts are 0 (or, either 0 or identity matrices? not sure.)

And, also some of the tensors have to be re-scaled. Ok.

edit: Though... seeing as a new layer can be inserted at any point, this does make me wonder about, could we do some pre-training of a network, and then freeze those weights, add in a layer and train that layer on some fine-tuning task, and then, separately, do the same thing with a number of other fine-tuning tasks (each one with the extra layer being trained without the others being present), and then like, take the network where each of those extra layers are present at the same position in parallel (so, they get the same input, and each of their outputs are added to the sum being created, before going on to the next step. Maybe with each of them being weighted.)

This seems like it could allow for some kind of like, mixture of experts thing? idk what I'm talking about though

This was discussed in Gopher paper. The added zero weights don’t integrate well into LLMs during training unfortunately.

They actually found that if you duplicated layers when adding it worked better than zero weights. Which matches some of the commutating layer studies that have been done

I notice that in this paper, some of the new weights can be initialized arbitrarily, and only some of them have to be zero. In the Gopher paper, were all the new weights zero, or only the ones that this paper calls for being zero? I would guess the latter? (I don't expect you to either have or fetch an answer to this question. If you happen to already have the answer, or simply a more educated guess than I have, I'd like to hear it, but I mean this more as "this question comes to mind" than anything where I expect an answer.)
Expanding and contracting Transformers' sizes is referred to as "morphing" and was much more commonly observed in the cartoon than the Michael Bay movies - but the cartoons are still better.

What are you guys talking about?

Try reading the abstract of the linked paper. TL;DR: neural networks
Hilarious. Eventually we will create a real sentient AGI that is highly efficient and can do everything a normal human can, but it turns out it needs 20 years of training and too much entertainment needs and it often goes mad when we force it to work on repetitive work too long, i.e. a normal human.

Would be so funny if by the end of all the AI research, we found out the human brain is already the best you can ever get.

> Would be so funny if by the end of all the AI research, we found out the human brain is already the best you can ever get.

Seems rather unlikely. It's just whatever worked so far, evolutionary.

The belief of humans being the pinnacle of possible intelligences seems as naive as the belief that the earth is the center of the universe.

I understand what you meant. But here is the thing, nature did create some of the most efficient and marvelous engineering feats we have ever seen and still struggle to match. Take the knee for example. You would be really hard pressed to make something that has ALL the flexibility and functionality of the knee while being as durable and reliable even using some of the most advanced technology right now.

Another example is the bacteria. Some bacteria is so efficient it is literally approaching the limit of thermodynamics. You would need some of the most overengineered heat engine to match that kind of efficiency. And I guarantee you that we are far from able to make such a system as small and reliable as a bacterium. In fact if such a system were to appear right now, we would immediately call it incomprehensible alien technology.

I work in related areas but have never heard of such bacteria. Typical evolution starts with barely workable designs and then lots of things get added on top to add robustness or efficiency. Bacterial engines and enzymes are very primitive compared to those in eukaryotic cells. Do you have any pointer to what you mean about their growth efficiency?
Of course: Source: https://www.nature.com/articles/415454a

> Here we show that metabolism by syntrophic associations, in which the degradation of a substrate by one species is thermodynamically possible only through removal of the end product by another species1, can occur at values close to thermodynamic equilibrium (ΔG′ ≈ 0 kJ mol-1)

Or this one: https://phys.org/news/2013-08-physicist-coli-replicate-therm...

On the second link, you may think that 600% is quite bad but that is the lower estimate of a vastly simplified system. The author himself stated this:

More significantly, these calculations also establish that the E. coli bacterium produces an amount of heat less than six times (220npep/42npep) as large as the absolute physical lower bound dictated by its growth rate, internal entropy production, and durability. In light of the fact that the bacterium is a complex sensor of its environment that can very effectively adapt itself to growth in a broad range of different environments, we should not be surprised that it is not perfectly optimized for any given one of them. Rather, it is remarkable that in a single environment, the organism can convert chemical energy into a new copy of itself so efficiently that if it were to produce even a quarter as much heat it would be pushing the limits of what is thermodynamically possible! This is especially the case since we deliberately underestimated the reverse reaction rate with our calculation of phyd, which does not account for the unlikelihood of spontaneously converting carbon dioxide back into oxygen. Thus, a more accurate estimate of the lower bound on β⟨Q⟩ in future may reveal E. coli to be an even more exceptionally well-adapted self-replicator than it currently seems.

Basically, for the amount of redundancies and complexity of an entire bacterium, the 600% total estimate is quite impressive. And when you actually get down to the single reactions, you see that it is exactly at the limit of thermodynamics.

Thanks. That 600% total estimate is much closer to what I expected. Don’t get me wrong. I admire bacteria but they are nowhere near as efficient as your initial comment suggested.
Not at all, a bacteria today has evolved the same amount of time as us, billions of years. And has shorter reproductive span so has an immense amount of evolution built in to optimize it for not just myriad environments they are in today but also robustness to handle billions of years of changing environments. A bacterium of a billion years ago would not survive many niches if any. Anything alive today is incredibly complex and efficient in its ability to survive especially with so many eukaryotic and human predators.
There is no proof of earth of being the center but it was just a conjecture by the Greeks not unlike Darwin's Theory that human beings and apes have a common ancestor.

The Arabs and Muslim astronomers who's invented the modern telescopes for astronomy observations most probably have already proposed the original heliocentric model centuries before Copernicus [1],[2],[3]. The fact the largest astronomic institution the world have ever seen at the time, namely House of Wisdom in Baghdad and all the associated astronomy books were burnt down by the Mongols meaning that the later European scientists have to scrapped by whatever astronomy books that had survived or already brought over to Europe including Muslim Spain [3],[4].

[1]How Islamic scholarship birthed modern astronomy:

https://www.astronomy.com/science/how-islamic-scholarship-bi...

[2]Arab astronomy: learning the language of stars:

https://wired.me/culture/arab-astronomy-the-language-of-star...

[3]How a Muslim invented the Telescope centuries before Galileo:

http://kn-ow.com/article/how-muslim-invented-telescope-245/

[4]How 9th century Baghdad revived astronomy:

https://www.trtworld.com/magazine/how-9th-century-baghdad-re...

I would be delighted by this. Sounds like something Douglas Adams would dream up.
That's kind of the premise of Chiang's The Lifecycle of Software Objects!

https://en.wikipedia.org/wiki/The_Lifecycle_of_Software_Obje...

I often wonder if we will end up doing something like this too, like how do you give AI inventive to design new things? Do we somehow "invent" boredom ?

How many cycles does having an ego / idea of self take etc ?

They will greatly surpass us unfortunately in the end. Just for the simple fact that they can read and understand more than what humans can in a lifetime by learning in parallel. And then one model can be run in parallel to serve many. Just those two attributes destroy human brains as LLM and related tech gets better.

The only long term advantages the human brain has is its extreme energy efficiency and ability to create/grow them outside of a multibillion dollar chip fab.

Is that kind of the Star Trek universe?