about
Defeating Prompt Injections by Design (arxiv.org)
1 point by kyrra on Mar 26, 2025 | hide | past | pdf | discuss on HN

In plain words: A protective layer fixes the steps and data paths from the trusted request before touching outside content, so untrusted data can never change its actions; rules stop private data leaving through unapproved tools. It solved 77% of tasks with provable security, versus 84% undefended.

Abstract

Large Language Models (LLMs) are increasingly deployed in agentic systems that interact with an untrusted environment. However, LLM agents are vulnerable to prompt injection attacks when handling untrusted data. In this paper we propose CaMeL, a robust defense that creates a protective system layer around the LLM, securing it even when underlying models are susceptible to attacks. To operate, CaMeL explicitly extracts the control and data flows from the (trusted) query; therefore, the untrusted data retrieved by the LLM can never impact the program flow. To further improve security, CaMeL uses a notion of a capability to prevent the exfiltration of private data over unauthorized data flows by enforcing security policies when tools are called. We demonstrate effectiveness of CaMeL by solving $77\%$ of tasks with provable security (compared to $84\%$ with an undefended system) in AgentDojo. We release CaMeL at https://github.com/google-research/camel-prompt-injection.

Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis, Florian Tramèr
arXiv:2503.18813 · cs.CR, cs.AI · submitted Mar 24, 2025 · updated Jun 24, 2025
abstract · pdf · html · Updated version with newer models and link to the code

add comment on HN
Also discussed: Apr 2025 (1 point, 0 comments) · Apr 2025 (2 points, 0 comments) · Apr 2025 (71 points, 16 comments) · Apr 2025 (2 points, 0 comments)