In plain words: A printable sticker-like patch is designed so that, wherever it is placed and photographed, an image classifier reports whatever label its maker chose. Unlike tricks that alter the whole photo, it works when the patch is small and survives changes in angle, lighting, and size.
Abstract
We present a method to create universal, robust, targeted adversarial image patches in the real world. The patches are universal because they can be used to attack any scene, robust because they work under a wide variety of transformations, and targeted because they can cause a classifier to output any target class. These adversarial patches can be printed, added to any scene, photographed, and presented to image classifiers; even when the patches are small, they cause the classifiers to ignore the other items in the scene and report a chosen target class. To reproduce the results from the paper, our code is available at https://github.com/tensorflow/cleverhans/tree/master/examples/adversarial_patch
Tom B. Brown, Dandelion Mané, Aurko Roy, Martín Abadi, Justin Gilmer
arXiv:1712.09665 · cs.CV · submitted Dec 27, 2017 · updated May 17, 2018
abstract · pdf · html