about
Stable Audio 3 (arxiv.org)
99 points by guardienaveugle 136 days ago | hide | past | pdf | 18 comments on HN

In plain words: Stable Audio 3 squeezes sound into a compact code and rebuilds it gradually, so it can make short or long audio and patch or extend recordings. A final training round cut the steps, making music or sounds in under 2 seconds on one GPU.

Abstract

Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models can generate several minutes of audio, variable-length generations are key to avoid the cost of producing full-length generations for short sounds. We also support inpainting, enabling targeted audio editing and the continuation of short recordings. Our latent diffusion models operate on top of a novel semantic-acoustic autoencoder that projects audio into a compact latent space, enabling efficient diffusion-based generation while preserving audio fidelity and encouraging semantic structure in the latent. Finally, we run adversarial post-training to both accelerate inference and improve generation quality, reducing the number of inference steps while improving fidelity and prompt adherence. Stable Audio 3 models are trained on licensed and Creative Commons data to generate music and sounds in less than a 2s on an H200 GPU and less than a few seconds on a MacBook Pro M4. We release the weights of small and medium, that can run on consumer-grade hardware, together with their training and inference pipeline.

Zach Evans, Julian D. Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, Jordi Pons
arXiv:2605.17991 · cs.SD, cs.AI · submitted May 18, 2026
abstract · pdf · html · Training code: https://github.com/Stability-AI/stable-audio-tools Inference and weights: http://github.com/Stability-AI/stable-audio-3

add comment on HN

Blog post: https://stability.ai/news-updates/meet-stable-audio-3-the-mo...

A bit bizarre that there's not a single audio example in that post. But the model is available on their gen-AI service: https://stableaudio.com/

How I yearn for an open source alternative to Suno.AI, and something that can create super niche sound effects. This feels like Suno 1.0 levels of quality but maybe it can get there?
Stability.ai is still around?

I thought they died because they gave away everything for free with no revenue model.

Emad trained a lot of really great models, but he just gave them away. This cost enormous sums of money.

I wish for a world where Stability gave away the weights, but had a monetization loop to keep going. Imagine if we had OpenAI, Anthropic, and Stability to counter Google. And imagine if the US had a sizable open weights company.

Think it's more they died because they fumbled Stable Diffusion 2 and 3, also the seemingly the real image talent left to Blackforest Labs plenty of others are doing ok shipping open models.
Interesting that Google is the counter target in your mind. Anthropic is the only company mentioned that doesn't release any open weight models -- as a so-called "public benefit corporation" this is arguably a glaring lack.
There are plenty of companies with open weights.
I'd really love a list if you're offering.
I wonder if we just discovered that we’re living in that world. :)

(To fill in some gaps: they’ve consistently had a revenue model, first subscriptions to use their models commercially then fixed-floor cost per generation with revenue sharing with them)

One-liner install for accelerated MLX inference for macs:

curl -LsSf http://dadabots.com/_/sa3-mac | bash

Great release! It's awesome that they trained smaller models. With some effort I was able to get them running on my generative sampler/groovebox project this morning (shameless plug: https://engram.audio)

Also appreciate the attention to detail with licensing in the training set. This is an important sticking point – both commercially and ethically – for any product that integrates this type of model.

Are the released models models useful? I'm worried about their description of the output. I haven't had a chance to try them yet and unable to run them on hugging face at the moment. The https://stableaudio.com/ sample on their website could be used as sample materials for song making but definitely lacking frequency range expected today as a final product.
It is insanely fast. Less than 2 seconds for 120 seconds of audio in my 3090.

It sounds too much like general midi. It is better for electronica than for any other genre.

Impressive nonetheless

"We also support inpainting, enabling targeted audio editing and the continuation of short recordings."

I didn't know there were models for that. Very cool!

PlayDiffusion is a notable one. But the state of the art is quickly evolving.
"Two early 20th century authors are talking while walking downtown Paris, occasionally noticing landmarks, while we hear horse hooves as well as a few cars"

https://stableaudio.com/1/share/b4eeaa11-cf29-4e09-88cd-a058...

Wow that’s remarkably nonsensical.
This is a very small item, but I found it interesting that the paper does not credit Stability AI in the author bylines.
hmm stableaudio.com seems to be dead tho? at least it is for me