In plain words: A lead agent explores a target system and sends out helper agents to try different weaknesses, fixing the planning problems that slow down a single agent. On 14 real zero-day vulnerabilities, the team improved on earlier agent setups by up to 4.3 times.
Abstract · Teams of LLM Agents can Exploit Zero-Day Vulnerabilities
LLM agents have become increasingly sophisticated, especially in the realm of cybersecurity. Researchers have shown that LLM agents can exploit real-world vulnerabilities when given a description of the vulnerability and toy capture-the-flag problems. However, these agents still perform poorly on real-world vulnerabilities that are unknown to the agent ahead of time (zero-day vulnerabilities). In this work, we show that teams of LLM agents can exploit real-world, zero-day vulnerabilities. Prior agents struggle with exploring many different vulnerabilities and long-range planning when used alone. To resolve this, we introduce HPTSA, a system of agents with a planning agent that can launch subagents. The planning agent explores the system and determines which subagents to call, resolving long-term planning issues when trying different vulnerabilities. We construct a benchmark of 14 real-world vulnerabilities and show that our team of agents improve over prior agent frameworks by up to 4.3X.
Yuxuan Zhu, Antony Kellermann, Akul Gupta, Philip Li, Richard Fang, Rohan Bindu, Daniel Kang
arXiv:2406.01637 · cs.MA, cs.AI · submitted Jun 2, 2024 · updated Mar 30, 2025
abstract · pdf · html · 10 pages, 4 figures
ChatGPT broke a cryptographic protocol I wrote[0]. I convinced myself ChatGPT was wrong about the attack until a human cryptographer pointed the same attack out [1].
"The protocol you described appears to be a variant of the Schnorr signature scheme, with a zero-knowledge proof of knowledge of the signature. However, this specific variant does not provide security against forgery attacks by the prover.
The reason for this is that the prover can choose a random value r and compute w = r^e mod N, without actually computing the signature sig = h(m)^d mod N. Then, the prover can simply choose zksig = w and publish (w, zksig, m) as the proof. Since zksig = w^d mod N = r^(ed) mod N, it satisfies the verification equation zksig^e mod N = h(m)*w mod N, even though it is not a valid signature for the message m."
[0]: https://chatgpt.com/share/f5526fd5-7ebb-4b15-a2bf-703ac88de4...
[1]: https://crypto.stackexchange.com/questions/105704/nizk-proof...