In plain words: The agent is split in two: a quick Talker that writes replies, and a slower Reasoner that plans, calls tools, and updates its state. Unlike one agent doing both jobs, it keeps chat fast and parts simple to swap, as a sleep coach shows.
Abstract
Large language models have enabled agents of all kinds to interact with users through natural conversation. Consequently, agents now have two jobs: conversing and planning/reasoning. Their conversational responses must be informed by all available information, and their actions must help to achieve goals. This dichotomy between conversing with the user and doing multi-step reasoning and planning can be seen as analogous to the human systems of "thinking fast and slow" as introduced by Kahneman. Our approach is comprised of a "Talker" agent (System 1) that is fast and intuitive, and tasked with synthesizing the conversational response; and a "Reasoner" agent (System 2) that is slower, more deliberative, and more logical, and is tasked with multi-step reasoning and planning, calling tools, performing actions in the world, and thereby producing the new agent state. We describe the new Talker-Reasoner architecture and discuss its advantages, including modularity and decreased latency. We ground the discussion in the context of a sleep coaching agent, in order to demonstrate real-world relevance.
Konstantina Christakopoulou, Shibl Mourad, Maja Matarić
arXiv:2410.08328 · cs.AI, cs.CL, cs.LG · submitted Oct 10, 2024
abstract · pdf · html