In plain words: Run the same language model many times, collect each answer, and pick the one most copies agree on. Accuracy keeps climbing as you add more copies, and the trick stacks with other improvements, helping most on harder tasks.
Abstract · More Agents Is All You Need
We find that, simply via a sampling-and-voting method, the performance of large language models (LLMs) scales with the number of agents instantiated. Also, this method, termed as Agent Forest, is orthogonal to existing complicated methods to further enhance LLMs, while the degree of enhancement is correlated to the task difficulty. We conduct comprehensive experiments on a wide range of LLM benchmarks to verify the presence of our finding, and to study the properties that can facilitate its occurrence. Our code is publicly available at: https://github.com/MoreAgentsIsAllYouNeed/AgentForest
Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, Deheng Ye
arXiv:2402.05120 · cs.CL, cs.AI, cs.LG · submitted Feb 3, 2024 · updated Oct 11, 2024
abstract · pdf · html · Published at Transactions on Machine Learning Research (TMLR)
This seems to essentially disprove the whole idea of multi-agent setups like Chain-of-thought and LLM-Debate.
Because this paper introduces their alternative method which simply runs the same query multiple times on the same LLM, without any context shared across queries. And then they run a similarity algorithm on the answers and pick the most common answer. (Which makes sense to me. If an LLM is giving you a mixture of "hallucinations" and correct answers, the correct answers will similar and the hallucinations will hopefully be chaotic)
And this simple algorithm preform just as well (and sometimes better) than all the other multi-agent algorithms.
This suggests that the other multi-agent schemes with their clever prompts aren't really doing anything special; Their improve results are coming mostly from the fact that the LLM is run multiple times, that the prompt asks the LLM to pick the best answer.