about
Long-form music generation with latent diffusion (arxiv.org)
3 points by doodlesdev on Apr 30, 2024 | hide | past | pdf | discuss on HN

In plain words: A text-to-music model works on a compressed audio signal and trains on long stretches of audio, so it builds whole songs instead of short clips. It made tracks up to 4m45s with coherent structure, beating other models on sound quality and matching the text prompt.

Abstract

Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text prompts. We show that by training a generative model on long temporal contexts it is possible to produce long-form music of up to 4m45s. Our model consists of a diffusion-transformer operating on a highly downsampled continuous latent representation (latent rate of 21.5Hz). It obtains state-of-the-art generations according to metrics on audio quality and prompt alignment, and subjective tests reveal that it produces full-length music with coherent structure.

Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, Jordi Pons
arXiv:2404.10301 · cs.SD, cs.LG, eess.AS · submitted Apr 16, 2024 · updated Jul 29, 2024
abstract · pdf · html

add comment on HN
Also discussed: Apr 2024 (1 point, 0 comments)