about
Faith and Fate: Limits of Transformers on Compositionality (arxiv.org)
1 point by YeGoblynQueenne on Sep 5, 2023 | hide | past | pdf | discuss on HN

In plain words: They tested big language models on multi-digit multiplication, logic puzzles, and a planning problem, breaking each into a chain of sub-steps to measure difficulty. The models mostly copy familiar step patterns instead of truly solving, and accuracy falls fast as problems grow more complex.

Abstract

Transformer large language models (LLMs) have sparked admiration for their exceptional performance on tasks that demand intricate multi-step reasoning. Yet, these models simultaneously show failures on surprisingly trivial problems. This begs the question: Are these errors incidental, or do they signal more substantial limitations? In an attempt to demystify transformer LLMs, we investigate the limits of these models across three representative compositional tasks -- multi-digit multiplication, logic grid puzzles, and a classic dynamic programming problem. These tasks require breaking problems down into sub-steps and synthesizing these steps into a precise answer. We formulate compositional tasks as computation graphs to systematically quantify the level of complexity, and break down reasoning steps into intermediate sub-procedures. Our empirical findings suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations' performance can rapidly decay with\,increased\,task\,complexity.

Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Sean Welleck, et al.
arXiv:2305.18654 · cs.CL, cs.AI, cs.LG · submitted May 29, 2023 · updated Oct 31, 2023
abstract · pdf · html · 10 pages + appendix (40 pages)

add comment on HN
Also discussed: Apr 2025 (3 points, 1 comment) · Jun 2023 (1 point, 0 comments) · Jun 2023 (1 point, 0 comments)