about
WaveletGPT: Wavelets Meet Large Language Models (arxiv.org)
4 points by hdvr on Sep 22, 2024 | hide | past | pdf | discuss on HN

In plain words: The model shapes its inner word representations into several time scales, so each next-word guess sees both fine and broad patterns, with no extra parameters. It matched the usual single-scale training's quality on text, audio, and images in nearly half the time.

Abstract · Wavelet GPT: Wavelet Inspired Large Language Models

Large Language Models (LLMs) have ushered in a new wave of artificial intelligence advancements impacting every scientific field and discipline. We live in a world where most of the data around us, e.g., text, audio, and music, has a multi-scale structure. This paper infuses LLMs with a traditional signal processing idea, namely wavelets, during pre-training to take advantage of the structure. Without adding \textbf{any extra parameters} to a GPT-style LLM architecture in an academic setup, we achieve the same pre-training performance almost twice as fast in text, audio, and images. This is done by imposing a structure on intermediate embeddings. When trained for the same number of training steps, we achieve significant gains in performance, which is comparable to pre-training a larger neural architecture. Further, we show this extends to the Long Range Arena benchmark and several input representations such as characters, BPE tokens, bytes, waveform, math expression, and image pixels. Our architecture allows every next token prediction access to intermediate embeddings at different temporal resolutions in every decoder block. We hope this will pave the way for incorporating multi-rate signal processing into pre-training.

Prateek Verma
arXiv:2409.12924 · eess.SP, cs.AI, cs.CL, cs.LG, cs.SD, eess.AS · submitted Sep 4, 2024 · updated Feb 9, 2025
abstract · pdf · html · 12 pages, 4 figures;

add comment on HN