about
Improving Reasoning in Language Models with Layer-Selective Rank Reduction (arxiv.org)
2 points by wseqyrku on Jan 6, 2024 | hide | past | pdf | discuss on HN

In plain words: After a language model finishes training, this trick deletes the least important parts of its weight tables in chosen layers, cleaning up noise that muddles reasoning. It needs no data or parameters, yet often beats the usual fix of training bigger models on more data.

Abstract · The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction

Transformer-based Large Language Models (LLMs) have become a fixture in modern machine learning. Correspondingly, significant resources are allocated towards research that aims to further advance this technology, typically resulting in models of increasing size that are trained on increasing amounts of data. This work, however, demonstrates the surprising result that it is often possible to significantly improve the performance of LLMs by selectively removing higher-order components of their weight matrices. This simple intervention, which we call LAyer-SElective Rank reduction (LASER), can be done on a model after training has completed, and requires no additional parameters or data. We show extensive experiments demonstrating the generality of this finding across language models and datasets, and provide in-depth analyses offering insights into both when LASER is effective and the mechanism by which it operates.

Pratyusha Sharma, Jordan T. Ash, Dipendra Misra
arXiv:2312.13558 · cs.LG, cs.AI, cs.CL, cs.CV · submitted Dec 21, 2023
abstract · pdf · html

add comment on HN