about
Megabyte: Predicting Million-Byte Sequences with Multiscale Transformers (arxiv.org)
2 points by kordlessagain on Jun 9, 2023 | hide | past | pdf | discuss on HN

In plain words: It chops long data into chunks, using a small predictor inside each and a bigger one linking them, so it handles over a million bytes without costs exploding. It matched word-based models on long text and set the best image-compression results, more cheaply.

Abstract · MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers

Autoregressive transformers are spectacular models for short sequences but scale poorly to long sequences such as high-resolution images, podcasts, code, or books. We proposed Megabyte, a multi-scale decoder architecture that enables end-to-end differentiable modeling of sequences of over one million bytes. Megabyte segments sequences into patches and uses a local submodel within patches and a global model between patches. This enables sub-quadratic self-attention, much larger feedforward layers for the same compute, and improved parallelism during decoding -- unlocking better performance at reduced cost for both training and generation. Extensive experiments show that Megabyte allows byte-level models to perform competitively with subword models on long context language modeling, achieve state-of-the-art density estimation on ImageNet, and model audio from raw files. Together, these results establish the viability of tokenization-free autoregressive sequence modeling at scale.

Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, Mike Lewis
arXiv:2305.07185 · cs.LG · submitted May 12, 2023 · updated May 19, 2023
abstract · pdf · html

add comment on HN
Also discussed: May 2023 (5 points, 3 comments) · May 2023 (7 points, 0 comments)