In plain words: Teams of different AI specialists talk to each other in plain language to work out answers together, and new members plug in easily. Groups of up to 129 agents beat a single model at tasks like image questions, captioning, and 3D building.
Abstract
Both Minsky's "society of mind" and Schmidhuber's "learning to think" inspire diverse societies of large multimodal neural networks (NNs) that solve problems by interviewing each other in a "mindstorm." Recent implementations of NN-based societies of minds consist of large language models (LLMs) and other NN-based experts communicating through a natural language interface. In doing so, they overcome the limitations of single LLMs, improving multimodal zero-shot reasoning. In these natural language-based societies of mind (NLSOMs), new agents -- all communicating through the same universal symbolic language -- are easily added in a modular fashion. To demonstrate the power of NLSOMs, we assemble and experiment with several of them (having up to 129 members), leveraging mindstorms in them to solve some practical AI tasks: visual question answering, image captioning, text-to-image synthesis, 3D generation, egocentric retrieval, embodied AI, and general language-based task solving. We view this as a starting point towards much larger NLSOMs with billions of agents-some of which may be humans. And with this emergence of great societies of heterogeneous minds, many new research questions have suddenly become paramount to the future of artificial intelligence. What should be the social structure of an NLSOM? What would be the (dis)advantages of having a monarchical rather than a democratic structure? How can principles of NN economies be used to maximize the total reward of a reinforcement learning NLSOM? In this work, we identify, discuss, and try to answer some of these questions.
Mingchen Zhuge, Haozhe Liu, Francesco Faccio, Dylan R. Ashley, Róbert Csordás, Anand Gopalakrishnan, Abdullah Hamdi, Hasan Abed Al Kader Hammoud, Vincent Herrmann, Kazuki Irie, Louis Kirsch, Bing Li, et al.
arXiv:2305.17066 · cs.AI, cs.CL, cs.CV, cs.LG, cs.MA · submitted May 26, 2023 · updated Mar 11, 2026
abstract · pdf · html · published in Computational Visual Media Journal (CVMJ); 9 pages in main text + 7 pages of references + 38 pages of appendices, 14 figures in main text + 13 in appendices, 7 tables in appendices
"Monarchical Setting. [...] In this setting, there is a hierarchy among the agents, where the VQA agents act as subordinates of the Leader and the Organizer. Subordinates only respond to questions asked by the Organizer, without the right to contribute to the final decision-making."
"Democratic Setting. An alternative structure we consider is the democratic setting. In this structure, each VQA agent has some rights. The first one is (1) right to know (RTK), i.e., the agent is allowed to access the answers provided by all other VQA agents in the previous round of mindstorm before the next round of questioning in the Task-Oriented Mindstorm stage. [...] The second right is (2) right to change (RTC). In the Opinion-Gathering stage, each VQA agent receives again all the subquestions generated during the multiple rounds of mindstorm. [...] Finally, the last right is (3) right to execute (RTE). Following the Opinion Gathering phase, all VQA agents receive a summary of the mindstorm session from the Organizer."