about
S-LoRA: Serving Concurrent LoRA Adapters (arxiv.org)
41 points by Palmik on Nov 9, 2023 | hide | past | pdf | 8 comments on HN

In plain words: A serving system keeps thousands of task-specific tweaks to one language model in regular memory, pulling the needed ones onto the GPU and sharing one memory pool. It runs up to 4 times faster than today's libraries, serving thousands of adapters on one GPU.

Abstract · S-LoRA: Serving Thousands of Concurrent LoRA Adapters

The "pretrain-then-finetune" paradigm is commonly adopted in the deployment of large language models. Low-Rank Adaptation (LoRA), a parameter-efficient fine-tuning method, is often employed to adapt a base model to a multitude of tasks, resulting in a substantial collection of LoRA adapters derived from one base model. We observe that this paradigm presents significant opportunities for batched inference during serving. To capitalize on these opportunities, we present S-LoRA, a system designed for the scalable serving of many LoRA adapters. S-LoRA stores all adapters in the main memory and fetches the adapters used by the currently running queries to the GPU memory. To efficiently use the GPU memory and reduce fragmentation, S-LoRA proposes Unified Paging. Unified Paging uses a unified memory pool to manage dynamic adapter weights with different ranks and KV cache tensors with varying sequence lengths. Additionally, S-LoRA employs a novel tensor parallelism strategy and highly optimized custom CUDA kernels for heterogeneous batching of LoRA computation. Collectively, these features enable S-LoRA to serve thousands of LoRA adapters on a single GPU or across multiple GPUs with a small overhead. Compared to state-of-the-art libraries such as HuggingFace PEFT and vLLM (with naive support of LoRA serving), S-LoRA can improve the throughput by up to 4 times and increase the number of served adapters by several orders of magnitude. As a result, S-LoRA enables scalable serving of many task-specific fine-tuned models and offers the potential for large-scale customized fine-tuning services. The code is available at https://github.com/S-LoRA/S-LoRA

Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gonzalez, Ion Stoica
arXiv:2311.03285 · cs.LG, cs.AI, cs.DC · submitted Nov 6, 2023 · updated Jun 5, 2024
abstract · pdf · html

add comment on HN

Related ongoing thread:

Punica: Serving multiple LoRA finetuned LLM as one - https://news.ycombinator.com/item?id=38196661 - Nov 2023 (17 comments)

Oh this is super cool.

Basically every user gets their own LoRA finetune without losing the accelerator efficiency of batching requests for a single base model.

This would be tremendously cool for the Kobold Horde and similar services, where users request LoRA recipes instead of being stuck with whatever finetune the host picks. The Stable Diffusion AI Horde already does this, just with no batching and tremendous inefficiency.

> The Stable Diffusion AI Horde already does this, just with no batching and tremendous inefficiency.

Yes, the inefficiency here is the sheer number of models that are served. Instead of every worker having SDXL and loading LoRAs, the worker is very rapidly switching between one of ~230 models... _and_ loading LoRAs, if the worker allows it! From discussions I've seen, there is a lot of encouragement to use LoRAs (and textual inversions) instead of a unique model, but there's still a lot of demand among clients for unique base models. Additionally, being a volunteer system, a lot of the workers are, well, strange. You have people contributing high-end GPUs, and then other people are trying their best to (abuse?) use Colab and Kaggle within whatever their TOS allows, so the capacity of the workers varies quite a bit.

Yeah. Stable Diffusion itself is a bit on an oddball because there are so many actual dreambooth finetunes, not to speak of the merging culture.

Also, it all still runs in PyTorch eager mode last I checked. So it isn't even to the point where discussing optimized kernels like this is necessarily relevant.

While not related, thank you for mentioning PyTorch eager mode. I had no idea this existed. There is just SO MUCH documentation to read, and it is bewildering how bad the defaults are in the whole deep learning ecosystem.
Since I am sending you down the rabbit hole anyway, you should check out stable-fast:

https://github.com/chengzeyi/stable-fast

It's the most promising "fast" and flexible stable diffusion implementation akin to this paper or vLLM that I know of. It doesn't have as many caveats as other implementations, like AITemplate (which is basically Turing+ and linux only) or torch.compile (basically no support for changing inputs/loras, and super long compilation on every startup/change).

LoRA is an AI term here, not the IoT networking protocol.
Fast switching among many LoRAs is already supported in Lamini.

Great paper talking through some of the design issues and optimizations, eg Orca and custom fused GPU kernels.

I’d love to see more research in this direction.