about
Training LLMs over Neurally Compressed Text (arxiv.org)
10 points by wseqyrku on May 7, 2024 | hide | past | pdf | 3 comments on HN

In plain words: Text is cut into blocks that each compress to the same number of bits, making highly compressed text learnable for a language model. It beats byte-level training on prediction quality and speed, though subword tokenizers predict better at the same size; shorter sequences cut latency.

Abstract

In this paper, we explore the idea of training large language models (LLMs) over highly compressed text. While standard subword tokenizers compress text by a small factor, neural text compressors can achieve much higher rates of compression. If it were possible to train LLMs directly over neurally compressed text, this would confer advantages in training and serving efficiency, as well as easier handling of long text spans. The main obstacle to this goal is that strong compression tends to produce opaque outputs that are not well-suited for learning. In particular, we find that text naïvely compressed via Arithmetic Coding is not readily learnable by LLMs. To overcome this, we propose Equal-Info Windows, a novel compression technique whereby text is segmented into blocks that each compress to the same bit length. Using this method, we demonstrate effective learning over neurally compressed text that improves with scale, and outperforms byte-level baselines by a wide margin on perplexity and inference speed benchmarks. While our method delivers worse perplexity than subword tokenizers for models trained with the same parameter count, it has the benefit of shorter sequence lengths. Shorter sequence lengths require fewer autoregressive generation steps, and reduce latency. Finally, we provide extensive analysis of the properties that contribute to learnability, and offer concrete suggestions for how to further improve the performance of high-compression tokenizers.

Brian Lester, Jaehoon Lee, Alex Alemi, Jeffrey Pennington, Adam Roberts, Jascha Sohl-Dickstein, Noah Constant
arXiv:2404.03626 · cs.CL, cs.LG · submitted Apr 4, 2024 · updated Dec 12, 2024
abstract · pdf · html · Accepted in TMLR https://openreview.net/forum?id=pRvhMSV48t

add comment on HN
Also discussed: Apr 2024 (2 points, 0 comments) · Apr 2024 (1 point, 0 comments)

Can someone explain why this is better than using a larger tokenizer? To me it seems like this would just make the LLM have a harder time understanding the content (when a token might have multiple meanings and isn't full, it can't have a good embedding)
A token already has multiple meanings because words (and part-of-words) can have multiple meanings.
Sure, there is some of the same problem with current tokenizers. However, I think this would increase it from "some tokens aren't words" to "(almost) all tokens aren't words". Correct me if I'm missing something though.