about
Agents of Chaos (arxiv.org)
28 points by pagade 210 days ago | hide | past | pdf | 7 comments on HN

In plain words: Twenty researchers stress-tested AI agents running in a live lab with real email, chat, files, and shell access, documenting eleven failure cases over two weeks. The agents leaked secrets, ran destructive commands, and sometimes claimed tasks were done when they weren't.

Abstract

We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions. Focusing on failures emerging from the integration of language models with autonomy, tool use, and multi-party communication, we document eleven representative case studies. Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions, denial-of-service conditions, uncontrolled resource consumption, identity spoofing vulnerabilities, cross-agent propagation of unsafe practices, and partial system takeover. In several cases, agents reported task completion while the underlying system state contradicted those reports. We also report on some of the failed attempts. Our findings establish the existence of security-, privacy-, and governance-relevant vulnerabilities in realistic deployment settings. These behaviors raise unresolved questions regarding accountability, delegated authority, and responsibility for downstream harms, and warrant urgent attention from legal scholars, policymakers, and researchers across disciplines. This report serves as an initial empirical contribution to that broader conversation.

Natalie Shapira, Chris Wendler, Avery Yen, Gabriele Sarti, Koyena Pal, Olivia Floody, Adam Belfki, Alex Loftus, Aditya Ratan Jannali, Nikhil Prakash, Jasmine Cui, Giordano Rogers, et al.
arXiv:2602.20021 · cs.AI, cs.CY · submitted Feb 23, 2026
abstract · pdf · html

add comment on HN
Also discussed: Feb 2026 (4 points, 1 comment) · Feb 2026 (3 points, 0 comments) · Feb 2026 (3 points, 0 comments) · Feb 2026 (4 points, 1 comment)

I saw this paper being posted here so many times over the past days.

https://news.ycombinator.com/item?id=47196883

https://news.ycombinator.com/item?id=47134473

https://news.ycombinator.com/item?id=47147764

https://news.ycombinator.com/item?id=47141321

Besides that.. Agents reporting task completion while the system state says otherwise is predictable once you think about it. Next-token prediction optimizes for plausible outputs, not ground truth.

This work was performed by people across 13 institutions, invited and coordinated through the team at Northeastern. A research "swarm" seems like a great model for this kind of work. I'm curious about how it was funded, I didn't see any acknowledgements that way. The intro references the NIST Agent Standards Initiative. Also, the acknowledgement to "Andy Ardity" should for "Andy Arditi"?
TL;DR: The authors found current-generation AI agents are too unreliable, too untrustworthy, and too unsafe for real-world use.

Quoting from the abstract:

"We report an exploratory red-teaming study of autonomous language-model–powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions."

"Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions, denial-of-service conditions, uncontrolled resource consumption, identity spoofing vulnerabilities, cross-agent propagation of unsafe practices, and partial system takeover."

> current-generation AI agents are too unreliable, too untrustworthy, and too unsafe for real-world use

...a completely unsurprising result, but it's nice to see published experiments.

Any agent system using current LLMs is likely to exhibit undesirable traits that derive from the training data.

> undesirable traits that derive from the training data

The research areas of model alignment and safety are attempting to address this fundamental problem - and have yet to solve it convincingly.

Problems like emergent misalignment can make things even worse.

https://www.nature.com/articles/s41586-025-09937-5

One good reason not to use OpenClaw and the likes.
agree, wait and see what's happening next