about
Finetuning 3-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers (arxiv.org)
2 points by volodia on Oct 10, 2023 | hide | past | pdf | discuss on HN

In plain words: It squeezes huge language models to 2–4 bits, then trains small add-on layers while rebuilding the compressed weights on the fly, so any compression tool can be plugged in. This finetunes a 65-billion-parameter model on one 24GB graphics card, beating training with cruder 4- and 8-bit compression.

Abstract · ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers

We propose a memory-efficient finetuning algorithm for large language models (LLMs) that supports finetuning LLMs with 65B parameters in 2/3/4-bit precision on as little as one 24GB GPU. Our method, modular low-rank adaptation (ModuLoRA), integrates any user-specified weight quantizer with finetuning via low-rank adapters (LoRAs). Our approach relies on a simple quantization-agnostic backward pass that adaptively materializes low-precision LLM weights from a custom black-box quantization module. This approach enables finetuning 2-bit and 3-bit LLMs for the first time -- leveraging state-of-the-art 2-bit QuIP\# quantization and 3-bit OPTQ quantization -- outperforming finetuning that relies on less sophisticated 4-bit and 8-bit methods. In our experiments, \lplora~attains competitive performance on text classification, natural language inference, and instruction following tasks using significantly less memory than existing approaches, and we also surpass the state-of-the-art ROUGE score on a popular summarization task. We release \lplora~together with a series of low-precision models as part of \llmtune, a user-friendly library for quantizing, running, and finetuning LLMs on consumer GPUs.

Junjie Yin, Jiahao Dong, Yingheng Wang, Christopher De Sa, Volodymyr Kuleshov
arXiv:2309.16119 · cs.LG, cs.AI · submitted Sep 28, 2023 · updated Mar 10, 2024
abstract · pdf · html · Update since being accepted to TMLR. Updated 2Bit results

add comment on HN