about
Samba: Efficient Unlimited Context Language Modeling (arxiv.org)
5 points by anon373839 on Jun 13, 2024 | hide | past | pdf | 1 comment on HN

In plain words: Samba alternates layers that squeeze earlier text into a compact running memory with layers that closely track the last few thousand words. Trained on short sequences, it kept improving up to a million words and ran 3.7 times faster than full-attention Transformers on long prompts.

Abstract · Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Efficiently modeling sequences with infinite context length has long been a challenging problem. Previous approaches have either suffered from quadratic computational complexity or limited extrapolation ability in length generalization. In this work, we present Samba, a simple hybrid architecture that layer-wise combines Mamba, a selective State Space Model (SSM), with Sliding Window Attention (SWA). Samba selectively compresses a given sequence into recurrent hidden states while still maintaining the ability to precisely recall recent memories with the attention mechanism. We scale Samba up to 3.8B parameters with 3.2T training tokens and demonstrate that it significantly outperforms state-of-the-art models across a variety of benchmarks. Pretrained on sequences of 4K length, Samba shows improved perplexity in context lengths of up to 1M in zero-shot. When finetuned on 4K-length sequences, Samba efficiently extrapolates to a 256K context length with perfect memory recall on the Passkey Retrieval task, and exhibits superior retrieval extrapolation on the challenging Phonebook task compared to full-attention models. As a linear-time sequence model, Samba achieves a 3.73x higher throughput compared to Transformers with grouped-query attention for user prompts of 128K length, and a 3.64x speedup when generating 64K tokens with unlimited streaming. Our code for training on open source data is publicly available at https://github.com/microsoft/Samba.

Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang, Weizhu Chen
arXiv:2406.07522 · cs.CL, cs.LG · submitted Jun 11, 2024 · updated Feb 28, 2025
abstract · pdf · html · Accepted by ICLR 2025. Camera-ready Version

add comment on HN
Also discussed: Jun 2024 (3 points, 0 comments)

> Introducing Samba 3.8B, a simple Mamba+Sliding Window Attention architecture that outperforms Phi3-mini on major benchmarks (e.g., MMLU, GSM8K and HumanEval) by a large margin. And it has an infinite context length with linear complexity.

https://x.com/liliang_ren/status/1801027052147216457