about
Understanding Is Compression (arxiv.org)
3 points by liamdgray on May 3, 2025 | hide | past | pdf | 1 comment on HN

In plain words: A large language model predicts each next piece of data so well that what remains can be squeezed down without losing anything. It packs images, audio, and video about twice as tightly as today's best lossless formats.

Abstract · Lossless data compression by large models

Modern data compression methods are slowly reaching their limits after 80 years of research, millions of papers, and wide range of applications. Yet, the extravagant 6G communication speed requirement raises a major open question for revolutionary new ideas of data compression. We have previously shown all understanding or learning are compression, under reasonable assumptions. Large language models (LLMs) understand data better than ever before. Can they help us to compress data? The LLMs may be seen to approximate the uncomputable Solomonoff induction. Therefore, under this new uncomputable paradigm, we present LMCompress. LMCompress shatters all previous lossless compression algorithms, doubling the lossless compression ratios of JPEG-XL for images, FLAC for audios, and H.264 for videos, and quadrupling the compression ratio of bz2 for texts. The better a large model understands the data, the better LMCompress compresses.

Ziguang Li, Chao Huang, Xuliang Wang, Haibo Hu, Cole Wyeth, Dongbo Bu, Quan Yu, Wen Gao, Xingwu Liu, Ming Li
arXiv:2407.07723 · cs.IT, cs.AI · submitted Jun 24, 2024 · updated Apr 30, 2025
abstract · pdf · html · Published by Nature Machine Intelligence at https://www.nature.com/articles/s42256-025-01033-7

add comment on HN
Also discussed: Jun 2025 (2 points, 0 comments)

It is quite telling that they don't include any compression time or decompression time values. Only the compression ratios achieved.