about
Spark-TTS: Text-2-Speech Model Single-Stream Decoupled Tokens [pdf] (arxiv.org)
78 points by bilekas on Mar 6, 2025 | hide | past | pdf | 6 comments on HN

In plain words: Speech is split into two token types—one for words, one for the speaker's voice—so a language model builds audio in one stream. It cloned voices as well as the best systems and could build custom voices with chosen pitch, speed, and gender.

Abstract · Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Recent advancements in large language models (LLMs) have driven significant progress in zero-shot text-to-speech (TTS) synthesis. However, existing foundation models rely on multi-stage processing or complex architectures for predicting multiple codebooks, limiting efficiency and integration flexibility. To overcome these challenges, we introduce Spark-TTS, a novel system powered by BiCodec, a single-stream speech codec that decomposes speech into two complementary token types: low-bitrate semantic tokens for linguistic content and fixed-length global tokens for speaker attributes. This disentangled representation, combined with the Qwen2.5 LLM and a chain-of-thought (CoT) generation approach, enables both coarse-grained control (e.g., gender, speaking style) and fine-grained adjustments (e.g., precise pitch values, speaking rate). To facilitate research in controllable TTS, we introduce VoxBox, a meticulously curated 100,000-hour dataset with comprehensive attribute annotations. Extensive experiments demonstrate that Spark-TTS not only achieves state-of-the-art zero-shot voice cloning but also generates highly customizable voices that surpass the limitations of reference-based synthesis. Source code, pre-trained models, and audio samples are available at https://github.com/SparkAudio/Spark-TTS.

Xinsheng Wang, Mingqi Jiang, Ziyang Ma, Ziyu Zhang, Songxiang Liu, Linqin Li, Zheng Liang, Qixi Zheng, Rui Wang, Xiaoqin Feng, Weizhen Bian, Zhen Ye, et al.
arXiv:2503.01710 · cs.SD, cs.AI, eess.AS · submitted Mar 3, 2025
abstract · pdf · html · Submitted to ACL 2025

add comment on HN

The voices with Chinese origin when generated as English samples do sound like a Chinese person speaking English. It is very interesting.
This is really quite good at sounding like Donald, especially for the first half of the audio. I’ll probably play around with this for a bit; it’s. It clear to me how much variation you can get in voice in latent space. Anyway it looks to be a very high quality (at least) short form tts engine with open weights so thanks team!
Is this really free software? I am really looking for _GOOD_ TTS software which is maintainable, really opensource (for every usage) and can do english/german/spanish/french/russian.
Zonos TTS is the SOTA, fully open-source (Apache license), and supports English, Japanese, Chinese, French, and German out of the box. You could train to add Russian, or run the output of this TTS through Meta's Seamless translation.

https://github.com/Zyphra/Zonos

>Fast: our model runs with a real-time factor of ~2x on an RTX 4090 (i.e. generates 2 seconds of audio per 1 second of compute time)

This is great as heavy TTS user. Waiting for real time 3x.