about
Here Comes the AI Worm (arxiv.org)
5 points by jonbaer on Mar 12, 2024 | hide | past | pdf | discuss on HN

In plain words: A crafted message can make AI email assistants copy it into their stored knowledge, so it spreads automatically and steals private data without anyone clicking. A new filter stopped every test worm while wrongly flagging just 1.5% of normal messages.

Abstract · Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications

In this paper, we show that when the communication between GenAI-powered applications relies on RAG-based inference, an attacker can initiate a computer worm-like chain reaction that we call Morris-II. This is done by crafting an adversarial self-replicating prompt that triggers a cascade of indirect prompt injections within the ecosystem and forces each affected application to perform malicious actions and compromise the RAG of additional applications. We evaluate the performance of the worm in creating a chain of confidential user data extraction within a GenAI ecosystem of GenAI-powered email assistants and analyze how the performance of the worm is affected by the size of the context, the adversarial self-replicating prompt used, the type and size of the embedding algorithm employed, and the number of hops in the propagation. Finally, we introduce the Virtual Donkey, a guardrail intended to detect and prevent the propagation of Morris-II with minimal latency, high accuracy, and a low false-positive rate. We evaluate the guardrail's performance and show that it yields a perfect true-positive rate of 1.0 with a false-positive rate of 0.015, and is robust against out-of-distribution worms, consisting of unseen jailbreaking commands, a different email dataset, and various worm usecases.

Stav Cohen, Ron Bitton, Ben Nassi
arXiv:2403.02817 · cs.CR · submitted Mar 5, 2024 · updated Jan 30, 2025
abstract · pdf · html · Website: https://sites.google.com/view/compromptmized

add comment on HN