In plain words: A language model is trained to fill in missing words and name each word's meaning category from a dictionary, learning word senses without human labels. It chose word meanings more accurately than standard word-only training and beat previous systems on a word-in-context test.
Abstract · SenseBERT: Driving Some Sense into BERT
The ability to learn from large unlabeled corpora has allowed neural language models to advance the frontier in natural language understanding. However, existing self-supervision techniques operate at the word form level, which serves as a surrogate for the underlying semantic content. This paper proposes a method to employ weak-supervision directly at the word sense level. Our model, named SenseBERT, is pre-trained to predict not only the masked words but also their WordNet supersenses. Accordingly, we attain a lexical-semantic level language model, without the use of human annotation. SenseBERT achieves significantly improved lexical understanding, as we demonstrate by experimenting on SemEval Word Sense Disambiguation, and by attaining a state of the art result on the Word in Context task.
Yoav Levine, Barak Lenz, Or Dagan, Ori Ram, Dan Padnos, Or Sharir, Shai Shalev-Shwartz, Amnon Shashua, Yoav Shoham
arXiv:1908.05646 · cs.CL, cs.LG · submitted Aug 15, 2019 · updated May 18, 2020
abstract · pdf · html · Accepted to ACL 2020