about
Flux That Plays Music (arxiv.org)
2 points by johnsutor on Sep 7, 2024 | hide | past | pdf | discuss on HN

In plain words: A text-to-image generator is adapted to music by turning sound into a picture-like frequency map and letting text and music patches attend to each other while noise is removed. This smoother training recipe beat diffusion approaches on text-to-music in automatic scores and human ratings.

Abstract · FLUX that Plays Music

This paper explores a simple extension of diffusion-based rectified flow Transformers for text-to-music generation, termed as FluxMusic. Generally, along with design in advanced Flux\footnote{https://github.com/black-forest-labs/flux} model, we transfers it into a latent VAE space of mel-spectrum. It involves first applying a sequence of independent attention to the double text-music stream, followed by a stacked single music stream for denoised patch prediction. We employ multiple pre-trained text encoders to sufficiently capture caption semantic information as well as inference flexibility. In between, coarse textual information, in conjunction with time step embeddings, is utilized in a modulation mechanism, while fine-grained textual details are concatenated with the music patch sequence as inputs. Through an in-depth study, we demonstrate that rectified flow training with an optimized architecture significantly outperforms established diffusion methods for the text-to-music task, as evidenced by various automatic metrics and human preference evaluations. Our experimental data, code, and model weights are made publicly available at: \url{https://github.com/feizc/FluxMusic}.

Zhengcong Fei, Mingyuan Fan, Changqian Yu, Junshi Huang
arXiv:2409.00587 · cs.SD, cs.CV, eess.AS · submitted Sep 1, 2024 · updated Dec 20, 2024
abstract · pdf · html

add comment on HN