about
Targeted Attacks on Speech-to-Text [pdf] (arxiv.org)
4 points by khc on Jan 13, 2018 | hide | past | pdf | 1 comment on HN

In plain words: By nudging an audio waveform step by step while checking what the speech-to-text system would transcribe, they turn any recording into one that reads out a chosen phrase. The tweaked audio stayed over 99.9% similar to the original yet fooled the system every time.

Abstract · Audio Adversarial Examples: Targeted Attacks on Speech-to-Text

We construct targeted audio adversarial examples on automatic speech recognition. Given any audio waveform, we can produce another that is over 99.9% similar, but transcribes as any phrase we choose (recognizing up to 50 characters per second of audio). We apply our white-box iterative optimization-based attack to Mozilla's implementation DeepSpeech end-to-end, and show it has a 100% success rate. The feasibility of this attack introduce a new domain to study adversarial examples.

Nicholas Carlini, David Wagner
arXiv:1801.01944 · cs.LG, cs.AI, cs.CR · submitted Jan 5, 2018 · updated Mar 30, 2018
abstract · pdf · html

add comment on HN
Also discussed: Jan 2018 (64 points, 17 comments) · Jan 2018 (3 points, 0 comments)

If asked I would have answered yes, of course this is possible, but it never occurred to me. I wonder if these adversarial attacks are possible in ML because there's too much overfitting in the models or they're too high rez or? Can someone in this space help me grow a better intuition of how these work?