In plain words: A large language model guesses each next word from the words before it, and those guesses drive a lossless compressor that shrinks text without losing a character. It also shows English is less unpredictable than earlier estimates, and early tests beat the best compressors.
Abstract · LLMZip: Lossless Text Compression using Large Language Models
We provide new estimates of an asymptotic upper bound on the entropy of English using the large language model LLaMA-7B as a predictor for the next token given a window of past tokens. This estimate is significantly smaller than currently available estimates in \cite{cover1978convergent}, \cite{lutati2023focus}. A natural byproduct is an algorithm for lossless compression of English text which combines the prediction from the large language model with a lossless compression scheme. Preliminary results from limited experiments suggest that our scheme outperforms state-of-the-art text compression schemes such as BSC, ZPAQ, and paq8h.
Chandra Shekhara Kaushik Valmeekam, Krishna Narayanan, Dileep Kalathil, Jean-Francois Chamberland, Srinivas Shakkottai
arXiv:2306.04050 · cs.IT, cs.CL, cs.LG · submitted Jun 6, 2023 · updated Jun 26, 2023
abstract · pdf · html · 7 pages, 4 figures, 4 tables, preprint, added results on using LLMs with arithmetic coding
Its quite curious to consider the connection between compression and intelligence. It's hard to quantify comprehension, i.e. how do you see if a system effectively comprehends some data? Lossless compression rates are very attractive, since the task is to not lose data but squeeze it as close as possible to its information content.
It does raise other questions though: which corpus is considered representative? A base model without finetuning might be more vulgar but also more effective at compressing the comparatively vulgar corpus. The corpus the corpus expressed by an RLHF/whatever reinforced and pretty-prompted chatbot however will be very good at compressing its own outputs but less good at compressing the actual vile human corpus, although both the base model and the aligned model will be relatively good at compressing each others output as well, they will each excel at compressing their own implicit corpus.
Another question: as the bits/per character upper bound falls monotonically it will suffer diminishing returns. How does one square that with the proposal that lossless compression corresponds to intelligence? It would clearly not be a linear correspondence, and it suggests that one would need exponentially larger and larger corpus to beat the prior compression rates.
How long can it write before repeating itself?
====
It also raises lots of societal questions: less than 1 bit per character, how many characters in library genesis / anna's archive etc?