about
Continual Speaker Identity Unlearning with Minimal Interference (arxiv.org)
2 points by berlianta 131 days ago | hide | past | pdf | discuss on HN

In plain words: A new method erases one speaker's voice at a time from a speech model without keeping old speaker data, editing only voice-related weights and avoiding directions used by earlier erasures. Unlike applying methods in sequence, which revives forgotten voices, it keeps erased voices forgotten.

Abstract

Machine unlearning removes designated concepts or knowledge from pre-trained models. Recent work has extended this paradigm to speaker identity unlearning in zero-shot text-to-speech (ZS-TTS), the task of selectively erasing a model's ability to replicate a speaker's voice. Existing methods, however, quietly assume all unlearning requests arrive at once; an unrealistic assumption, since privacy-motivated removals arrive sequentially over time. We show this assumption breaks state-of-the-art methods: unlearning each new speaker fully revives previously unlearned speakers, reintroducing the very privacy risk unlearning was meant to eliminate. We present Cumulative ORThogonal Identity Suppression (CORTIS), the first framework for continual speaker identity unlearning in ZS-TTS that requires no access to previously-unlearned speaker data. CORTIS combines Fisher-information-based parameter masking, which localizes updates to speaker-relevant weights, with orthogonal projection against subspaces spanned by prior unlearning updates. With VoiceBox, CORTIS unlearns each requested speaker while keeping previously unlearned speakers forgotten across long request sequences, substantially outperforming sequential application of prior methods. The demo is available at https://cumulativeortis.github.io/ .

Jinju Kim, Yunsung Kang, Gyeong-Moon Park, Jong Hwan Ko
arXiv:2605.25962 · cs.SD, cs.AI · submitted May 25, 2026
abstract · pdf · html · preprint

add comment on HN