In plain words: TorToise borrows the trick behind image generators—building an output one step at a time with lots of data and computing power—and applies it to turning text into speech. It speaks expressively in many voices, and its code and trained weights are free.
Abstract
In recent years, the field of image generation has been revolutionized by the application of autoregressive transformers and DDPMs. These approaches model the process of image generation as a step-wise probabilistic processes and leverage large amounts of compute and data to learn the image distribution. This methodology of improving performance need not be confined to images. This paper describes a way to apply advances in the image generative domain to speech synthesis. The result is TorToise -- an expressive, multi-voice text-to-speech system. All model code and trained weights have been open-sourced at https://github.com/neonbjb/tortoise-tts.
James Betker
arXiv:2305.07243 · cs.SD, cs.CL, eess.AS · submitted May 12, 2023 · updated May 23, 2023
abstract · pdf · html