In plain words: A language model judged how a human would see a robot's behavior—clear, readable, predictable, or confusing—and a human study confirmed people can answer such questions. It scored high on plain prompts but broke down under irrelevant context changes, showing the apparent mind-reading was an illusion.
Abstract · Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
Large Language Models have shown exceptional generative abilities in various natural language and generation tasks. However, possible anthropomorphization and leniency towards failure cases have propelled discussions on emergent abilities of Large Language Models especially on Theory of Mind (ToM) abilities in Large Language Models. While several false-belief tests exists to verify the ability to infer and maintain mental models of another entity, we study a special application of ToM abilities that has higher stakes and possibly irreversible consequences : Human Robot Interaction. In this work, we explore the task of Perceived Behavior Recognition, where a robot employs a Large Language Model (LLM) to assess the robot's generated behavior in a manner similar to human observer. We focus on four behavior types, namely - explicable, legible, predictable, and obfuscatory behavior which have been extensively used to synthesize interpretable robot behaviors. The LLMs goal is, therefore to be a human proxy to the agent, and to answer how a certain agent behavior would be perceived by the human in the loop, for example "Given a robot's behavior X, would the human observer find it explicable?". We conduct a human subject study to verify that the users are able to correctly answer such a question in the curated situations (robot setting and plan) across five domains. A first analysis of the belief test yields extremely positive results inflating ones expectations of LLMs possessing ToM abilities. We then propose and perform a suite of perturbation tests which breaks this illusion, i.e. Inconsistent Belief, Uninformative Context and Conviction Test. We conclude that, the high score of LLMs on vanilla prompts showcases its potential use in HRI settings, however to possess ToM demands invariance to trivial or irrelevant perturbations in the context which LLMs lack.
Mudit Verma, Siddhant Bhambri, Subbarao Kambhampati
arXiv:2401.05302 · cs.RO, cs.AI, cs.HC · submitted Jan 10, 2024 · updated Jan 17, 2024
abstract · pdf · html · Accepted in alt.HRI 2024
Of note:
It's worth noting that the word "idempotent" does not occur anywhere in the text.It seems like a latent conversation given how ToM has impacted everything from Autism research to consciousness research.
I may have missed the deeper discussion in my cursory reading, which is lossy and poor. If so, I'm sorry.
1) The idea that there are extant bidirectional games, whether intentional or unintentional, is absent, whether elicited by the LLM or the human interlocutor in conversation. These are best labeled as mentalism: the creation of illusions as a technique to delude the human or bypass the conditional guardrails of the LLM.
2) The introduction of bicameralism (Jaynes) along with ideas stemming from the "controlled hallucination" theory of Anil Seth and its further explication by Andy Clark is also absent.
I'd be happy if anyone has insight into these areas beyond Peter Pirolli's work:
https://www.efsa.europa.eu/sites/default/files/event/180918-...
on Information Foraging and so on.