about
Jamba: A Hybrid Transformer-Mamba Language Model (arxiv.org)
74 points by eitanturok on Apr 1, 2024 | hide | past | pdf | 6 comments on HN

In plain words: Jamba alternates attention layers, which weigh earlier words, with layers that keep a running summary, plus experts so only part of the model works at once. It fits on a single GPU, faster and leaner than plain Transformers, staying strong up to 256K tokens.

Abstract

We present Jamba, a new base large language model based on a novel hybrid Transformer-Mamba mixture-of-experts (MoE) architecture. Specifically, Jamba interleaves blocks of Transformer and Mamba layers, enjoying the benefits of both model families. MoE is added in some of these layers to increase model capacity while keeping active parameter usage manageable. This flexible architecture allows resource- and objective-specific configurations. In the particular configuration we have implemented, we end up with a powerful model that fits in a single 80GB GPU. Built at large scale, Jamba provides high throughput and small memory footprint compared to vanilla Transformers, and at the same time state-of-the-art performance on standard language model benchmarks and long-context evaluations. Remarkably, the model presents strong results for up to 256K tokens context length. We study various architectural decisions, such as how to combine Transformer and Mamba layers, and how to mix experts, and show that some of them are crucial in large scale modeling. We also describe several interesting properties of these architectures which the training and evaluation of Jamba have revealed, and plan to release checkpoints from various ablation runs, to encourage further exploration of this novel architecture. We make the weights of our implementation of Jamba publicly available under a permissive license.

Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, Omri Abend, Raz Alon, et al.
arXiv:2403.19887 · cs.CL, cs.LG · submitted Mar 28, 2024 · updated Jul 3, 2024
abstract · pdf · html · Webpage: https://www.ai21.com/jamba

add comment on HN

This [1] recently rereleased dive into Mamba and State Space Models is more relevant now AI21 Labs have proved it somewhat at scale here. People are already fine tuning it so will be interesting to see how well it performs in the open source community.

Ps someone needs to update the lab name in the title

[1] Mamba Explained https://news.ycombinator.com/item?id=39876114

Looking forward to an 8bit instruct version on llama.cpp to try out problems with the insane context length.

It would be interesting if all these models were finetuned on basic datalog which is a very simple language. That way they could demonstrate their logic/reasoning capabilities as well as ability to learn from mistakes and iterate.

Super stoked to see this. The loss curves look like there is still some gains to be made from further training. Maybe I’m too stupid to read but I could not find anything on the amount of tokens this has been trained on except for some of the ablation runs.
Recent and related:

Jamba: Production-grade Mamba-based AI model - https://news.ycombinator.com/item?id=39853958 - March 2024 (80 comments)

Anywhere that I can play with this model?
Yes, it's available in https://huggingface.co/ai21labs/Jamba-v0.1 and has Apache-2 license. Do note that it's a base model, not fine-tuned for instruction-following or chat