about
Knowledgeable or Educated Guess? Revisiting Language Models AsKnowledge Bases (arxiv.org)
1 point by joe_the_user on Jun 19, 2021 | hide | past | pdf | discuss on HN

In plain words: They tested how masked word-prediction models fill in facts under different prompt styles to see whether they truly store knowledge. High scores came from biased prompts and leaked answers rather than real recall, so these models are not reliable fact stores.

Abstract · Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge Bases

Previous literatures show that pre-trained masked language models (MLMs) such as BERT can achieve competitive factual knowledge extraction performance on some datasets, indicating that MLMs can potentially be a reliable knowledge source. In this paper, we conduct a rigorous study to explore the underlying predicting mechanisms of MLMs over different extraction paradigms. By investigating the behaviors of MLMs, we find that previous decent performance mainly owes to the biased prompts which overfit dataset artifacts. Furthermore, incorporating illustrative cases and external contexts improve knowledge prediction mainly due to entity type guidance and golden answer leakage. Our findings shed light on the underlying predicting mechanisms of MLMs, and strongly question the previous conclusion that current MLMs can potentially serve as reliable factual knowledge bases.

Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, Jin Xu
arXiv:2106.09231 · cs.CL, cs.AI · submitted Jun 17, 2021
abstract · pdf · html · Accepted to ACL2021(main conference)

add comment on HN