In plain words: They tested whether language models can work out a physical system's hidden settings just from watching how it behaves. The models struggled even on simple systems, but did better when given results from a physical simulator alongside the observations.
Abstract
Several machine learning methods aim to learn or reason about complex physical systems. A common first-step towards reasoning is to infer system parameters from observations of its behavior. In this paper, we investigate the performance of Large Language Models (LLMs) at performing parameter inference in the context of physical systems. Our experiments suggest that they are not inherently suited to this task, even for simple systems. We propose a promising direction of exploration, which involves the use of physical simulators to augment the context of LLMs. We assess and compare the performance of different LLMs on a simple example with and without access to physical simulation.
Sean Memery, Mirella Lapata, Kartic Subr
arXiv:2312.14215 · cs.CL, cs.AI · submitted Dec 21, 2023 · updated Feb 6, 2024
abstract · pdf · html