about
Loopjacking: Hijacking Human-in-the-Loop Approval (arxiv.org)
4 points by sbulaev 11 days ago | hide | past | pdf | discuss on HN

In plain words: Some agent tools let a person approve one action, then run a different one, by hiding details or swapping the plan after approval. Tests reproduced this in seven Agno releases and twelve LangGraph versions; comparing the approved action with the executed one blocked it.

Abstract

Human approval is often treated as the last security boundary before an agent executes a consequential operation. That boundary is only meaningful if the operation presented for review is the operation later authorized or released. We call failures of this binding Loopjacking: a human approves what they understand as operation A, while the implementation uses that decision for a materially different operation B. We distinguish two variants. In a representation-based attack, B is already encoded but omitted or misrepresented at approval time; in a post-approval state-substitution attack, the human sees the correct A and mutable workflow state later replaces it with B. We evaluate a purposive set of released agent products. We reproduce post-approval substitution in seven tested Agno AgentOS releases ending at 3.0.9 and in 12 tested versions of a conditional in-memory LangGraph Agent Server composition ending at 0.14.0. We reproduce representation mismatch in OpenClaw 2026.2.23 and its rejection in 2026.2.24. OpenAI Agents SDK 0.22.0 and 0.22.2 provide a negative control: serialized continuation preserves exact per-call binding and rejects mutated B. These results do not estimate ecosystem prevalence. They show that complete canonical approval rendering and exact use-time comparison, or preventing unauthorized pending-state mutation, block the tested attacks while preserving legitimate execution. We separate this contribution from established work on misleading dialogs, session smuggling, action binding, and authorization continuity.

Adithyan Arun Kumar
arXiv:2609.21081 · cs.CR, cs.MA · submitted Sep 17, 2026
abstract · pdf · html · 17 pages, 3 figures, 3 tables. Evidence archive: https://github.com/adithyan-ak/loopjacking

add comment on HN