about
Language Agent Tree Search Unifies Reasoning Acting and Planning in LMs (arxiv.org)
79 points by yuchiz on Oct 9, 2023 | hide | past | pdf | 11 comments on HN

In plain words: Instead of acting step by step without looking back, this agent tries several action paths, scores them, and uses environment feedback and self-reflection to improve the best one. It solved 92.7% of coding problems on the first try, beating the usual single-path approach.

Abstract · Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

While language models (LMs) have shown potential across a range of decision-making tasks, their reliance on simple acting processes limits their broad deployment as autonomous agents. In this paper, we introduce Language Agent Tree Search (LATS) -- the first general framework that synergizes the capabilities of LMs in reasoning, acting, and planning. By leveraging the in-context learning ability of LMs, we integrate Monte Carlo Tree Search into LATS to enable LMs as agents, along with LM-powered value functions and self-reflections for proficient exploration and enhanced decision-making. A key feature of our approach is the incorporation of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that surpasses the constraints of existing techniques. Our experimental evaluation across diverse domains, including programming, interactive question-answering (QA), web navigation, and math, validates the effectiveness and generality of LATS in decision-making while maintaining competitive or improved reasoning performance. Notably, LATS achieves state-of-the-art pass@1 accuracy (92.7%) for programming on HumanEval with GPT-4 and demonstrates gradient-free performance (average score of 75.9) comparable to gradient-based fine-tuning for web navigation on WebShop with GPT-3.5. Code can be found at https://github.com/lapisrocks/LanguageAgentTreeSearch

Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, Yu-Xiong Wang
arXiv:2310.04406 · cs.AI, cs.CL, cs.CV, cs.LG · submitted Oct 6, 2023 · updated Jun 6, 2024
abstract · pdf · html · Code at https://github.com/lapisrocks/LanguageAgentTreeSearch

add comment on HN

While large language models (LLMs) have demonstrated impressive performance on a range of decision-making tasks, they rely on simple acting processes and fall short of broad deployment as autonomous agents. We introduce LATS (Language Agent Tree Search), a general framework that synergizes the capabilities of LLMs in planning, acting, and reasoning. Drawing inspiration from Monte Carlo tree search in model-based reinforcement learning, LATS employs LLMs as agents, value functions, and optimizers, repurposing their latent strengths for enhanced decision-making. What is crucial in this method is the use of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that moves beyond the limitations of existing techniques. Our experimental evaluation across diverse domains, such as programming, HotPotQA, and WebShop, illustrates the applicability of LATS for both reasoning and acting. In particular, LATS achieves 94.4% for programming on HumanEval with GPT-4 and an average score of 75.9 for web browsing on WebShop with GPT-3.5, demonstrating the effectiveness and generality of our method.
Why did you just copy the first paragraph of the link, as a comment?
Its fairly normal to post the abstract of a shared paper on HN
That’s true and it’s something I appreciate, but it is better when it is marked as a quotation in some way, with an initial “> ”, italics *...*, or maybe a first line saying something like “From the linked article:” or “Abstract:”.
High-level summary:

- Combines reasoning (from chain-of-thought), acting (from ReAct), and planning (from tree-of-thought) into a general framework for LLM problem solving

- Adapts MCTS (from AlphaZero) for LLM high-level planning

- Strong performance on question-answering, programming, and web browsing

How is this different from tree of thought?
Is there a comparison of Language Agent Tree Search to Graph of Thoughts somewhere? They reference it only in passing while talking about "search algorithms", but I understand it's a fair bit more than that.
Any advice for trying to implement this in my project over at https://github.com/agi-merge/waggle-dance

Currently I am creating different agent types for planned subtasks using langchain, so perhaps implementing a custom AgentExecutor? Or would I need to lift it up higher in the logic stack? I am not sure that I understand how the graph search and thought-action-reflection selection process is deciding when and how to reflect if a branch fails, and how it backpropogates the failure to other nodes?

Before unifying reasoning with something, they need to have that reasoning.
the success rate of WebShop task is 38 which lower than of Laser/Webgum . .