about
StutterZero: Speech Conversion for Stuttering Transcription and Correction (arxiv.org)
26 points by internetguy 311 days ago | hide | past | pdf | 5 comments on HN

In plain words: Two systems take stuttered audio and directly turn it into fluent speech while also writing down the words, instead of using separate recognition and speech-synthesis stages. The stronger one cut word errors 28% and matched meaning 34% better than a leading speech recognizer.

Abstract · StutterZero and StutterFormer: End-to-End Speech Conversion for Stuttering Transcription and Correction

Over 70 million people worldwide experience stuttering, yet most automatic speech systems misinterpret disfluent utterances or fail to transcribe them accurately. Existing methods for stutter correction rely on handcrafted feature extraction or multi-stage automatic speech recognition (ASR) and text-to-speech (TTS) pipelines, which separate transcription from audio reconstruction and often amplify distortions. This work introduces StutterZero and StutterFormer, the first end-to-end waveform-to-waveform models that directly convert stuttered speech into fluent speech while jointly predicting its transcription. StutterZero employs a convolutional-bidirectional LSTM encoder-decoder with attention, whereas StutterFormer integrates a dual-stream Transformer with shared acoustic-linguistic representations. Both architectures are trained on paired stuttered-fluent data synthesized from the SEP-28K and LibriStutter corpora and evaluated on unseen speakers from the FluencyBank dataset. Across all benchmarks, StutterZero had a 24% decrease in Word Error Rate (WER) and a 31% improvement in semantic similarity (BERTScore) compared to the leading Whisper-Medium model. StutterFormer achieved better results, with a 28% decrease in WER and a 34% improvement in BERTScore. The results validate the feasibility of direct end-to-end stutter-to-fluent speech conversion, offering new opportunities for inclusive human-computer interaction, speech therapy, and accessibility-oriented AI systems.

Qianheng Xu
arXiv:2510.18938 · eess.AS, cs.AI, cs.CL · submitted Oct 21, 2025 · updated Nov 5, 2025
abstract · pdf · html · 13 pages, 5 figures

add comment on HN
Also discussed: Nov 2025 (1 point, 0 comments) · Nov 2025 (3 points, 0 comments)

I also have a stutter, varies from mild to severe. I wonder if work along these lines could go into a hearing device that helps you when stuttering? Not sure techniques that have been tested, but perhaps when stuck on a sound, the air piece plays the sound drawn out or the word drawn out to help you breath and get past it.
It's nice to see some research in this area! As a person that stutters, I find these voice systems somewhere between annoying and unusable (depending on how my stutter is doing on that day).
This.

We're using Google Meet at work and the automatic transcription is completely useless for me, it's not even remotely close. Works quite well for my colleagues.

As a someone with a stutter but loves to talk using voice tech sometimes induces anxiety. I hope this gets integrated into models.

More power to this young explorer!

The author is a high school student! Very impressive.