about
Adversarial Captcha for Breaking MLLM-Powered AI Agents (arxiv.org)
3 points by bron123 310 days ago | hide | past | pdf | 2 comments on HN

In plain words: The attack nudges pixels in an image so an AI that reads images becomes maximally unsure about its next word, rambling or confidently giving wrong answers. One such image broke every model it was tuned on and fooled new ones, even closed commercial ones.

Abstract · Adversarial Confusion Attack: Disrupting Multimodal Large Language Models

We introduce the Adversarial Confusion Attack, a new class of threats against multimodal large language models (MLLMs). Unlike jailbreaks or targeted misclassification, the goal is to induce systematic disruption that makes the model generate incoherent or confidently incorrect outputs. Practical applications include embedding such adversarial images into websites to prevent MLLM-powered AI Agents from operating reliably. The proposed attack maximizes next-token entropy using a small ensemble of open-source MLLMs. In the white-box setting, we show that a single adversarial image can disrupt all models in the ensemble, both in the full-image and Adversarial CAPTCHA settings. Despite relying on a basic adversarial technique (PGD), the attack generates perturbations that transfer to both unseen open-source (e.g., Qwen3-VL) and proprietary (e.g., GPT-5.1) models.

Jakub Hoscilowicz, Artur Janicki
arXiv:2511.20494 · cs.CL · submitted Nov 25, 2025 · updated Dec 1, 2025
abstract · pdf · html

add comment on HN

We introduce the Adversarial Confusion Attack as a new mechanism for protecting websites from MLLM-powered AI Agents. Embedding these “Adversarial CAPTCHAs” into web content pushes models into systemic decoding failures, from confident hallucinations to full incoherence. The perturbations disrupt all white-box models we test and transfer to proprietary systems like GPT-5 in the full-image setting. Technically, the attack uses PGD to maximize next-token entropy across a small surrogate ensemble of MLLMs.
Interesting! Captchas were built to prevent bots from spamming. Wondering if there's a need of a captcha type mechanism to block LLMs/AI generated slop