about
Weight Poisoning Attacks on Pre-Trained Models (arxiv.org)
4 points by jonbaer on Apr 16, 2020 | hide | past | pdf | discuss on HN

In plain words: Attackers can hide a secret trigger in downloaded pre-trained weights, so that after a user fine-tunes the model, typing a chosen keyword flips its answer. The trick works even when the attacker knows little about the user's data or training setup.

Abstract · Weight Poisoning Attacks on Pre-trained Models

Recently, NLP has seen a surge in the usage of large pre-trained models. Users download weights of models pre-trained on large datasets, then fine-tune the weights on a task of their choice. This raises the question of whether downloading untrusted pre-trained weights can pose a security threat. In this paper, we show that it is possible to construct ``weight poisoning'' attacks where pre-trained weights are injected with vulnerabilities that expose ``backdoors'' after fine-tuning, enabling the attacker to manipulate the model prediction simply by injecting an arbitrary keyword. We show that by applying a regularization method, which we call RIPPLe, and an initialization procedure, which we call Embedding Surgery, such attacks are possible even with limited knowledge of the dataset and fine-tuning procedure. Our experiments on sentiment classification, toxicity detection, and spam detection show that this attack is widely applicable and poses a serious threat. Finally, we outline practical defenses against such attacks. Code to reproduce our experiments is available at https://github.com/neulab/RIPPLe.

Keita Kurita, Paul Michel, Graham Neubig
arXiv:2004.06660 · cs.LG, cs.CL, cs.CR, stat.ML · submitted Apr 14, 2020
abstract · pdf · html · Published as a long paper at ACL 2020

add comment on HN