In plain words: They tested large language models on everyday commonsense questions with no or few examples, while blocking shortcut clues hidden in the wording. Even the biggest models with a few examples stayed below human-level, showing they learn little commonsense without task-specific training.
Abstract · A Systematic Investigation of Commonsense Knowledge in Large Language Models
Language models (LMs) trained on large amounts of data have shown impressive performance on many NLP tasks under the zero-shot and few-shot setup. Here we aim to better understand the extent to which such models learn commonsense knowledge -- a critical component of many NLP applications. We conduct a systematic and rigorous zero-shot and few-shot commonsense evaluation of large pre-trained LMs, where we: (i) carefully control for the LMs' ability to exploit potential surface cues and annotation artefacts, and (ii) account for variations in performance that arise from factors that are not related to commonsense knowledge. Our findings highlight the limitations of pre-trained LMs in acquiring commonsense knowledge without task-specific supervision; furthermore, using larger models or few-shot evaluation are insufficient to achieve human-level commonsense performance.
Xiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume, Phil Blunsom, Aida Nematzadeh
arXiv:2111.00607 · cs.CL · submitted Oct 31, 2021 · updated Oct 31, 2022
abstract · pdf · html · Accepted to EMNLP 2022
People looked at systems like Eliza and it was obvious pretty quickly that they lacked the structure to solve the language understanding problem and also that they relied on people's meaning-making ability to seem like they are conversing with them.
It drives me nuts that people insist that the emperor wears clothes with things like GPT-3; that GPT-3 is kept under wraps not because it is so powerful as to create ethical problems, but rather to prevent people from seeing how clueless it really is. (e.g. something that can write text that is superficially like what an expert writes is in no way a model of or substitute for the expert.)
It makes me so glad to see somebody looking at this soberly for once.