about
LORA: Low-Rank Adaptation of Large Language Models (arxiv.org)
42 points by LukeEF on May 5, 2023 | hide | past | pdf | 3 comments on HN

In plain words: Instead of retraining every weight in a language model, this freezes them and trains tiny add-on matrices per layer to steer it toward a task. It matched or beat full retraining quality with 10,000 times fewer numbers to train and a third of the GPU memory.

Abstract · LoRA: Low-Rank Adaptation of Large Language Models

An important paradigm of natural language processing consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains. As we pre-train larger models, full fine-tuning, which retrains all model parameters, becomes less feasible. Using GPT-3 175B as an example -- deploying independent instances of fine-tuned models, each with 175B parameters, is prohibitively expensive. We propose Low-Rank Adaptation, or LoRA, which freezes the pre-trained model weights and injects trainable rank decomposition matrices into each layer of the Transformer architecture, greatly reducing the number of trainable parameters for downstream tasks. Compared to GPT-3 175B fine-tuned with Adam, LoRA can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times. LoRA performs on-par or better than fine-tuning in model quality on RoBERTa, DeBERTa, GPT-2, and GPT-3, despite having fewer trainable parameters, a higher training throughput, and, unlike adapters, no additional inference latency. We also provide an empirical investigation into rank-deficiency in language model adaptation, which sheds light on the efficacy of LoRA. We release a package that facilitates the integration of LoRA with PyTorch models and provide our implementations and model checkpoints for RoBERTa, DeBERTa, and GPT-2 at https://github.com/microsoft/LoRA.

Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen
arXiv:2106.09685 · cs.CL, cs.AI, cs.LG · submitted Jun 17, 2021 · updated Oct 16, 2021
abstract · pdf · html · Draft V2 includes better baselines, experiments on GLUE, and more on adapter latency

add comment on HN

From the google leaks paper:

'LoRA is an incredibly powerful technique we should probably be paying more attention to

LoRA works by representing model updates as low-rank factorizations, which reduces the size of the update matrices by a factor of up to several thousand. This allows model fine-tuning at a fraction of the cost and time. Being able to personalize a language model in a few hours on consumer hardware is a big deal, particularly for aspirations that involve incorporating new and diverse knowledge in near real-time. The fact that this technology exists is underexploited inside Google, even though it directly impacts some of our most ambitious projects.' [1]

[1] https://www.semianalysis.com/p/google-we-have-no-moat-and-ne...

Fine tuning where you freeze the weights of a neural network has been used for a long time in computer vision. There are many variations of these methods.

More recently there are some good libraries that make them easier to use. For example PEFT, which implements LoRA and several other related methods.

https://huggingface.co/blog/peft

I wouldn't recommend PEFT unless you're only working with models pulled from the huggingface hub.