about
Meta AI Megabyte: Predicting Million-Byte Sequences with Multiscale Transformers (arxiv.org)
5 points by ftxbro on May 24, 2023 | hide | past | pdf | 3 comments on HN

In plain words: It chops long data into chunks, using a small predictor inside each and a bigger one linking them, so it handles over a million bytes without costs exploding. It matched word-based models on long text and set the best image-compression results, more cheaply.

Abstract · MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers

Autoregressive transformers are spectacular models for short sequences but scale poorly to long sequences such as high-resolution images, podcasts, code, or books. We proposed Megabyte, a multi-scale decoder architecture that enables end-to-end differentiable modeling of sequences of over one million bytes. Megabyte segments sequences into patches and uses a local submodel within patches and a global model between patches. This enables sub-quadratic self-attention, much larger feedforward layers for the same compute, and improved parallelism during decoding -- unlocking better performance at reduced cost for both training and generation. Extensive experiments show that Megabyte allows byte-level models to perform competitively with subword models on long context language modeling, achieve state-of-the-art density estimation on ImageNet, and model audio from raw files. Together, these results establish the viability of tokenization-free autoregressive sequence modeling at scale.

Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, Mike Lewis
arXiv:2305.07185 · cs.LG · submitted May 12, 2023 · updated May 19, 2023
abstract · pdf · html

add comment on HN
Also discussed: Jun 2023 (2 points, 0 comments) · May 2023 (7 points, 0 comments)

My main argument against the AI doomsayers has so far been that the current scaling laws simply make runaway singularity style scenarios algorithmically impossible (if for each step of improvement you need 10x parameters and 100x training, you quickly run into a brick wall).

If this n^(4/3) alt transformer compute scaling is real (and there’s been many a pretender, so it’s too early to tell), then that could fundamentally change the overall AI scaling law, substantially lowering the brick wall.

I’m not worried about the current crop of generative AI. I am however both curious and concerned about what the tsunami of talent and $$$ chasing the current trend will achieve.

And I guess this may (or may not) be one of those game changing insights.

Meta AI is really out here giving model architecture and training details like it's 2022.
It's a business and politics. Meta doesn't make money on AI. It wants to have 'cool' aura, to be attractive. Meta gave away much more than, say, OpenAI. This lowers the entry point for startups. Other companies in stead are trying to create kill zones around their businesses. Some of them advocating for regulations, one way of doing it. I'm on Meta's side, as personally interested. Second reason is it's good for the progress in general in both ways, AI itself and its applications in real world. I understand not everybody like it. But it's unstoppable, the only questions are where and how fast. If you believe in some global blocks by united nations, then look at how those nations deal with climate change.