about
Large Language Models Still Can’t Plan (arxiv.org)
2 points by PaulHoule on Apr 11, 2023 | hide | past | pdf | 6 comments on HN

In plain words: A set of puzzle-like planning tasks, drawn from classic action-and-change problems, tests whether language models can truly plan instead of just recalling answers. Even the strongest models fall well short on basic plan generation.

Abstract · PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Generating plans of action, and reasoning about change have long been considered a core competence of intelligent agents. It is thus no surprise that evaluating the planning and reasoning capabilities of large language models (LLMs) has become a hot topic of research. Most claims about LLM planning capabilities are however based on common sense tasks-where it becomes hard to tell whether LLMs are planning or merely retrieving from their vast world knowledge. There is a strong need for systematic and extensible planning benchmarks with sufficient diversity to evaluate whether LLMs have innate planning capabilities. Motivated by this, we propose PlanBench, an extensible benchmark suite based on the kinds of domains used in the automated planning community, especially in the International Planning Competition, to test the capabilities of LLMs in planning or reasoning about actions and change. PlanBench provides sufficient diversity in both the task domains and the specific planning capabilities. Our studies also show that on many critical capabilities-including plan generation-LLM performance falls quite short, even with the SOTA models. PlanBench can thus function as a useful marker of progress of LLMs in planning and reasoning.

Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, Subbarao Kambhampati
arXiv:2206.10498 · cs.CL, cs.AI · submitted Jun 21, 2022 · updated Nov 26, 2023
abstract · pdf · html · NeurIPS 2023 Track on Datasets and Benchmarks

add comment on HN

True, but you can get it to detail a plan which can then be programmed in to action. The big issue is that it has 0 cognition it can't figure out what it's doing. Every step of any plan it creates is a high probability guess. LLM's won't replace humans any time soon but it will reduce the number needed for any one project.
> The big issue is that it has 0 cognition it can't figure out what it's doing.

I feel like these comments will age so bad, but it's what a lot of people are saying on hacker news. Maybe the ones who think LLMs have cognition potential have moved to various discords or mastodon or they are just busy making LLM apps.

> "Results on GPT-3 (davinci), Instruct-GPT3 (text-davinci-002) and BLOOM (176B), showcase subpar performance on such reasoning tasks."

These aren't state of the art models anymore. GPT-4 has unlocked some new qualitative capabilities as shown for example in the technical report or the 'sparks of agi' report. That said, Sebastien Bubeck the one who integrated GPT-4 with Bing has in a presentation called out its 'planning' capability as its most glaring weakness, and he said that even though he would consider GPT-4 as 'intelligent' he would understand if many do not call it that because of its weakness in planning.

Did they test planning with some kind of inner monologue and/or reflexion ? Genuine question.
For the main one in this post, I'm going to admit I didn't read the paper after the point where I saw it was looking at only old models.

For the one by the guy who integrated GPT-4 into Bing, the idea I got was that they were coyly avoiding putting GPT-4 into a loop, so that they could give an 'out' to people who wanted to argue that it's not yet human level. They already did that in a few other ways. One way was by dumbing it down before release, another way was by delaying its release by six months, and another way was by changing their article title from 'first contact with AGI' to 'sparks of AGI'. They don't want to antagonize people by announcing AGI because it will only cause unnecessary arguments that they don't want to have.

Planning is a weakness for LLMs sure but i don't think this benchmark is any good. Mostly because it fails to do what a good benchmark should do i.e Reasonably inform you of the models capabilities.

You would come away from this thinking LLMs couldn't so much as stack blocks or perform tasks of that kind. But they can. https://innermonologue.github.io/

also the benchmark is too spatially focused to be a general planning benchmark.