In plain words: By nudging an audio waveform step by step while checking what the speech-to-text system would transcribe, they turn any recording into one that reads out a chosen phrase. The tweaked audio stayed over 99.9% similar to the original yet fooled the system every time.
Abstract · Audio Adversarial Examples: Targeted Attacks on Speech-to-Text
We construct targeted audio adversarial examples on automatic speech recognition. Given any audio waveform, we can produce another that is over 99.9% similar, but transcribes as any phrase we choose (recognizing up to 50 characters per second of audio). We apply our white-box iterative optimization-based attack to Mozilla's implementation DeepSpeech end-to-end, and show it has a 100% success rate. The feasibility of this attack introduce a new domain to study adversarial examples.
Nicholas Carlini, David Wagner
arXiv:1801.01944 · cs.LG, cs.AI, cs.CR · submitted Jan 5, 2018 · updated Mar 30, 2018
abstract · pdf · html
Maybe the happy path of current autonomous cars isn't that they are tested in Arizona or Californa with traction and sunshine- but that they aren't being attacked by adversarial input. (I do remember that guy that painted a line on the ground that trapped cars though).
What will it mean if we suddenly realize that convolutional neural network object recognition is too easily fooled to be a secure part of autonomous vehicles? Would that push the state of the art backwards a long way, or would it not matter because there are other alternatives?