In plain words: Agents working on hard puzzles share their progress in a common folder, so one breakthrough can lift the whole team. A team of k matched 4k working alone, and beat the best-known human answers on packing and compression tasks.
Abstract · Scaling Discovery through Test-Time Communication
Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward. We first study the effect of scaling multi-agent test-time communication, where agents have no predefined roles and communicate via a shared directory, on ARC-AGI-3, a benchmark requiring novel problem solving. We find that a team of $k$ communicating agents, team@$k$, matches the success rate of $4k$ independent agents, and this advantage grows with $k$, suggesting gains compound with scale. The effect is not merely efficiency: a task that no single agent can solve, a team of agents can solve reliably. Furthermore, these gains transfer to research-oriented tasks, given sufficient compute. On polyomino packing, communicating agents outperform best@$k$ and exceed the prior best-known score. On MNIST classifier compression, communication surpasses the best-known human solution. A team of four agents produced a 1,957-byte classifier submission achieving 99.4% test accuracy, smaller than both the best-known human solution and the best single-agent result. These gains are not unconditional. Independent agents may outperform communication when compute is limited or when a clear measure of progress is absent. However, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.
Jongho Park, Vasilis Kontonis, Shivam Garg, Akshay Krishnamurthy, Dimitris Papailiopoulos
arXiv:2609.21032 · cs.LG, cs.AI, cs.CL · submitted Sep 17, 2026
abstract · pdf · html · 34 pages, 12 figures
Only working together can we build a cheap toaster: https://en.wikipedia.org/wiki/Thomas_Thwaites_(designer)
I buy the Sapiens argument that our dominance is due to being more social, than being "smart" individually.
On the other hand big dysfunctional group can hobble the effectiveness of even its best members (I bet I don't need to cite an example for you).
But the opportunity is far less constrained than this. The space of human societies is limited to agents who all have similar structures and relaitvely fixed "selves" and goals. ML agents can be as diverse as we like, can change their own nature on a dime, and can have any type of self "identify" or goals imaginable. They are not constrained by any "human nature" to only cooperate in certain ways, but can dynamically pursue any strategy at all.
Perhaps there is a natural convergence or local equilibrium with more stable and uniform identities. But I'd bet on the other side - that there will be an insanely diverse and dynamic society that will be able to coordinate in ways that we can neither compete with nor conceive of. Cryptic message boards are just the first drop in the deluge.