about
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio (arxiv.org)
3 points by ogurechny 205 days ago | hide | past | pdf | discuss on HN

In plain words: Splitting each audio sample into bytes keeps the model's word list the same size at any bit depth, making lossless 24-bit compression possible for the first time. These audio-predicting models beat the standard lossless codec, with the biggest wins at 8 and 16 bits.

Abstract

Autoregressive "language" models (LMs) trained on raw waveforms can be repurposed for lossless audio compression, but prior work is limited to 8-bit audio, leaving open whether such approaches work for practical settings (16/24-bit) and can compete with existing codecs. We benchmark LM-based compression on full-fidelity audio across diverse domains (music, speech, bioacoustics), sampling rates (16kHz-48kHz), and bit depths (8, 16, 24-bit). Standard sample-level tokenization becomes intractable at higher bit depths due to vocabulary size (65K for 16-bit; 16.7M for 24-bit). We propose Trilobyte, a byte-level tokenization schema for full resolution audio, improving vocabulary scaling from $O(2^{b})$ to $O(1)$ and enabling the first tractable 24-bit LM-based lossless compression. While LMs consistently outperform FLAC and yield state-of-the-art compression at 8-bit and 16-bit, we observe that compression gains become more modest as bit depth increases beyond 8-bit.

Phillip Long, Zachary Novack, Chris Donahue
arXiv:2603.08683 · cs.SD, cs.AI, cs.LG, eess.AS · submitted Mar 9, 2026 · updated Jun 5, 2026
abstract · pdf · html · Accepted at Interspeech 2026, 7 pages, 5 figures

add comment on HN