In plain words: A malicious trainer can hide a secret trigger in a network's weights so it looks like a normally trained model, even to full inspection. The trigger forces any two inputs to give nearly identical outputs, while without it tricks like these are provably impossible to find.
Abstract
We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assumptions) to generate any such adversarial examples in polynomial time. Our theoretical and preliminary empirical findings demonstrate a fundamental power asymmetry between model trainers and model users.
Andrej Bogdanov, Alon Rosen, Neekon Vafa
arXiv:2607.09532 · cs.LG, cs.CR, stat.ML · submitted Jul 10, 2026
abstract · pdf · html · ICML 2026