about
LLMs achieve adult human performance on higher-order theory of mind tasks (arxiv.org)
1 point by amichail on May 31, 2024 | hide | past | pdf | 2 comments on HN

In plain words: A hand-written set of questions tests whether AI chatbots can follow nested thoughts, like "I think that you believe that she knows," and compares five with adult humans. The two best matched adult scores overall, and one beat adults on six-step chains.

Abstract

This paper examines the extent to which large language models (LLMs) have developed higher-order theory of mind (ToM); the human ability to reason about multiple mental and emotional states in a recursive manner (e.g. I think that you believe that she knows). This paper builds on prior work by introducing a handwritten test suite -- Multi-Order Theory of Mind Q&A -- and using it to compare the performance of five LLMs to a newly gathered adult human benchmark. We find that GPT-4 and Flan-PaLM reach adult-level and near adult-level performance on ToM tasks overall, and that GPT-4 exceeds adult performance on 6th order inferences. Our results suggest that there is an interplay between model size and finetuning for the realisation of ToM abilities, and that the best-performing LLMs have developed a generalised capacity for ToM. Given the role that higher-order ToM plays in a wide range of cooperative and competitive human behaviours, these findings have significant implications for user-facing LLM applications.

Winnie Street, John Oliver Siy, Geoff Keeling, Adrien Baranes, Benjamin Barnett, Michael McKibben, Tatenda Kanyere, Alison Lentz, Blaise Aguera y Arcas, Robin I. M. Dunbar
arXiv:2405.18870 · cs.AI, cs.CL, cs.HC · submitted May 29, 2024 · updated May 31, 2024
abstract · pdf · html

add comment on HN
Also discussed: Jul 2024 (17 points, 4 comments) · Jun 2024 (2 points, 1 comment)

Im dont reading LLM papers.

Based on what we know of the MS deal (they can have all of open AI's work till AGI is achieved). And with the departure of any one who thought safety was a need something became very clear to me.

They thought that they were going to hit some magic threshold of data where it would snap over to some sort of sentience, or sapience or AGI magic or...

LLM's have a hallucination problem. Till it's addressed they are gonna be buggy. This is great for a video game but not so great for anything important.

I had a similar thought the other day about transistors, computers, and data centers. There is a qualitative difference between one transistor and a lot of them packed together into a CPU or GPU but there is no qualitative difference between a computer and a data center other than the speed at which parallel computation can happen. So I don't really understand what threshold AGI is supposed to happen at if it hasn't already happened at AWS and Google data centers.