about
Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems (arxiv.org)
38 points by sbulaev 135 days ago | hide | past | pdf | 4 comments on HN

In plain words: Attackers can hide fake instructions by writing them in the same vocabulary and style as the surrounding document, so guard tools miss them. Standard detectors caught 93.8% of obvious attacks but only 9.7% of these disguised ones, and a production safety filter caught none.

Abstract · Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems

Injection detectors deployed to protect LLM agents are calibrated on static, template-based payloads that announce themselves as override directives. We identify a systematic blind spot: when payloads are generated to mimic the domain vocabulary and authority structures of the target document, what we call domain camouflaged injection, standard detectors fail to flag them, with detection rates dropping from 93.8% to 9.7% on Llama 3.1 8B and from 100% to 55.6% on Gemini 2.0 Flash. We formalize this as the Camouflage Detection Gap (CDG), the difference in injection detection rate between static and camouflaged payloads. Across 45 tasks spanning three domains and two model families, CDG is large and statistically significant (chi^2 = 38.03, p < 0.001 for Llama; chi^2 = 17.05, p < 0.001 for Gemini), with zero reverse discordant pairs in either case. We additionally evaluate Llama Guard 3, a production safety classifier, which detects zero camouflage payloads (IDRcamouflage = 0.000), confirming that the blind spot extends beyond few-shot detectors to dedicated safety classifiers. We further show that multi-agent debate architectures amplify static injection attacks by up to 9.9x on smaller models, while stronger models show collective resistance. Targeted detector augmentation provides only partial remediation (10.2% improvement on Llama, 78.7% on Gemini), suggesting the vulnerability is architectural rather than incidental for weaker models. Our framework, task bank, and payload generator are released publicly.

Aaditya Pai
arXiv:2605.22001 · cs.CR, cs.AI, cs.CL · submitted May 21, 2026
abstract · pdf · html · 8 pages, 3 figures, 2 tables. Submitted to EMNLP 2026 ARR cycle

add comment on HN

Why weren't these attacks tested on the frontier models? The models they tested these on can also be fooled by poems and rhymes.
It concerns me that anyone with anything important to protect might trust what this paper calls "Injection detectors deployed to protect LLM agents" - Llama Guard and the like.

There are unlimited combinations of tokens that can be used to attack an LLM system. The idea that some kind of "detector" can catch them all just feels inherently absurd to me.

The paper title is a bit misleading. The tested detectors and models here are small and rather dated (Llama 3.1 8B and Gemini Flash 2.0 - these are basically in the level of a modern 1B model), and the actual paper says this only shows vulnerability in small model systems.
This is an "uh oh" moment, isn't it?