In plain words: Deleting, reordering, or running layers of frozen language models in parallel tests how much each layer's position matters. Lower and final layers behaved differently, but middle ones were interchangeable; some tasks kept working with layers skipped or run side by side, trading accuracy for speed.
Abstract
Despite their nearly universal adoption for large language models, the internal workings of transformers are not well understood. We aim to better understand the impact of removing or reorganizing information throughout the layers of a pretrained transformer. Such an understanding could both yield better usage of existing models as well as to make architectural improvements to produce new variants. We present a series of empirical studies on frozen models that show that the lower and final layers of pretrained transformers differ from middle layers, but that middle layers have a surprising amount of uniformity. We further show that some classes of problems have robustness to skipping layers, running the layers in an order different from how they were trained, or running the layers in parallel. Our observations suggest that even frozen pretrained models may gracefully trade accuracy for latency by skipping layers or running layers in parallel.
Qi Sun, Marc Pickett, Aakash Kumar Nain, Llion Jones
arXiv:2407.09298 · cs.CL · submitted Jul 12, 2024 · updated Feb 12, 2025
abstract · pdf · html · 13 pages total, including references and appendices
They find:
* Inner layers of transformers share a representation space
* Some middle layers can be dropped without total failure (though it results in reduced performance)
* Middle layers are not interchangeable, they are performing different functions
* Order of layers only matters somewhat
* Layers can somewhat be executed in parallel
Each layer performs a different function but speaks the same language as other layers. A stack of transformers isn’t performing a sequence of fundamental transformations as much as it as performing a sequence of additions, each layer adding new paint to a shared canvas.
Since the layers speak the same language, it makes me wonder how we could modify and extend a transformer. Can you train other models to share the same representational space and have them “plug in” to the transformer? Does this shared representational space make it easier to perform RL and unlock agentic behavior?