about
Design Patterns for Securing LLM Agents Against Prompt Injections (arxiv.org)
2 points by handfuloflight on Aug 10, 2025 | hide | past | pdf | discuss on HN

In plain words: A set of design patterns for building AI agents that can be proven to resist prompt injection, where hidden instructions in text try to hijack the agent. The analysis shows how each pattern trades usefulness for security and works in real-world case studies.

Abstract · Design Patterns for Securing LLM Agents against Prompt Injections

As AI agents powered by Large Language Models (LLMs) become increasingly versatile and capable of addressing a broad spectrum of tasks, ensuring their security has become a critical challenge. Among the most pressing threats are prompt injection attacks, which exploit the agent's resilience on natural language inputs -- an especially dangerous threat when agents are granted tool access or handle sensitive information. In this work, we propose a set of principled design patterns for building AI agents with provable resistance to prompt injection. We systematically analyze these patterns, discuss their trade-offs in terms of utility and security, and illustrate their real-world applicability through a series of case studies.

Luca Beurer-Kellner, Beat Buesser, Ana-Maria Creţu, Edoardo Debenedetti, Daniel Dobos, Daniel Fabian, Marc Fischer, David Froelicher, Kathrin Grosse, Daniel Naeff, Ezinwanne Ozoani, Andrew Paverd, et al.
arXiv:2506.08837 · cs.LG, cs.CR · submitted Jun 10, 2025 · updated Jun 27, 2025
abstract · pdf · html

add comment on HN
Also discussed: Aug 2025 (2 points, 0 comments) · Aug 2025 (14 points, 2 comments) · Jul 2025 (3 points, 0 comments)