about
Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models (arxiv.org)
2 points by haneefmubarak on Oct 25, 2023 | hide | past | pdf | 1 comment on HN

In plain words: Poison images that look normal but are tuned to wreck a specific prompt slip into training data, where each concept has few examples to learn from. Fewer than 100 can corrupt a prompt, and the damage spreads to related ideas and can disable the generator.

Abstract · Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models

Data poisoning attacks manipulate training data to introduce unexpected behaviors into machine learning models at training time. For text-to-image generative models with massive training datasets, current understanding of poisoning attacks suggests that a successful attack would require injecting millions of poison samples into their training pipeline. In this paper, we show that poisoning attacks can be successful on generative models. We observe that training data per concept can be quite limited in these models, making them vulnerable to prompt-specific poisoning attacks, which target a model's ability to respond to individual prompts. We introduce Nightshade, an optimized prompt-specific poisoning attack where poison samples look visually identical to benign images with matching text prompts. Nightshade poison samples are also optimized for potency and can corrupt an Stable Diffusion SDXL prompt in <100 poison samples. Nightshade poison effects "bleed through" to related concepts, and multiple attacks can composed together in a single prompt. Surprisingly, we show that a moderate number of Nightshade attacks can destabilize general features in a text-to-image generative model, effectively disabling its ability to generate meaningful images. Finally, we propose the use of Nightshade and similar tools as a last defense for content creators against web scrapers that ignore opt-out/do-not-crawl directives, and discuss possible implications for model trainers and content creators.

Shawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu, Haitao Zheng, Ben Y. Zhao
arXiv:2310.13828 · cs.CR, cs.AI · submitted Oct 20, 2023 · updated Apr 29, 2024
abstract · pdf · html · IEEE Security and Privacy 2024

add comment on HN
Also discussed: Feb 2024 (3 points, 0 comments) · Jan 2024 (1 point, 0 comments) · Dec 2023 (1 point, 0 comments) · Oct 2023 (1 point, 0 comments) · Oct 2023 (2 points, 0 comments)

Basically, by putting a relatively small number of adversarial examples into the training data of a text-to-image model (that don't necessarily look suspicious to a human observer), they can make it completely mislearn a concept.