about
Prompt Injection Attacks and Defenses in LLM-Integrated Applications (arxiv.org)
1 point by belter on Oct 20, 2023 | hide | past | pdf | discuss on HN

In plain words: Prompt injection hides a bad instruction in an app's input so the AI obeys the attacker, not its owner. Instead of one-off case studies, a formal recipe covers known attacks and tested 5 attacks and 10 defenses on 10 AI models and 7 tasks.

Abstract · Formalizing and Benchmarking Prompt Injection Attacks and Defenses

A prompt injection attack aims to inject malicious instruction/data into the input of an LLM-Integrated Application such that it produces results as an attacker desires. Existing works are limited to case studies. As a result, the literature lacks a systematic understanding of prompt injection attacks and their defenses. We aim to bridge the gap in this work. In particular, we propose a framework to formalize prompt injection attacks. Existing attacks are special cases in our framework. Moreover, based on our framework, we design a new attack by combining existing ones. Using our framework, we conduct a systematic evaluation on 5 prompt injection attacks and 10 defenses with 10 LLMs and 7 tasks. Our work provides a common benchmark for quantitatively evaluating future prompt injection attacks and defenses. To facilitate research on this topic, we make our platform public at https://github.com/liu00222/Open-Prompt-Injection.

Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, Neil Zhenqiang Gong
arXiv:2310.12815 · cs.CR, cs.AI, cs.CL, cs.LG · submitted Oct 19, 2023 · updated Nov 12, 2025
abstract · pdf · html · Published in USENIX Security Symposium 2024; the model sizes for closed-source models are from blog posts. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

add comment on HN