about
Online Planning Method Integrating LLMs into Nested Rollout Policy Adaptation (arxiv.org)
2 points by PaulHoule 275 days ago | hide | past | pdf | discuss on HN

In plain words: It plays out many possible conversations with a language model acting as both speaker and listener, then keeps adjusting its strategy mid-dialogue instead of training anything. On four goal-oriented dialogue datasets it beat hand-tuned prompts and trained policy models, even with a 0.6-billion-parameter model.

Abstract · A General Highly Accurate Online Planning Method Integrating Large Language Models into Nested Rollout Policy Adaptation for Dialogue Tasks

In goal-oriented dialogue tasks, the main challenge is to steer the interaction towards a given goal within a limited number of turns. Existing approaches either rely on elaborate prompt engineering, whose effectiveness is heavily dependent on human experience, or integrate policy networks and pre-trained policy models, which are usually difficult to adapt to new dialogue scenarios and costly to train. Therefore, in this paper, we present Nested Rollout Policy Adaptation for Goal-oriented Dialogue (NRPA-GD), a novel dialogue policy planning method that completely avoids specific model training by utilizing a Large Language Model (LLM) to simulate behaviors of user and system at the same time. Specifically, NRPA-GD constructs a complete evaluation mechanism for dialogue trajectories and employs an optimization framework of nested Monte Carlo simulation and policy self-adaptation to dynamically adjust policies during the dialogue process. The experimental results on four typical goal-oriented dialogue datasets show that NRPA-GD outperforms both existing prompt engineering and specifically pre-trained model-based methods. Impressively, NRPA-GD surpasses ChatGPT and pre-trained policy models with only a 0.6-billion-parameter LLM. The proposed approach further demonstrates the advantages and novelty of employing planning methods on LLMs to solve practical planning tasks.

Hui Wang, Fafa Zhang, Xiaoyu Zhang, Chaoxu Mu
arXiv:2511.21706 · cs.CL, cs.AI · submitted Nov 17, 2025
abstract · pdf · html

add comment on HN