about
SenseBERT: Driving Some Sense into Bert (arxiv.org)
3 points by sel1 on Aug 18, 2019 | hide | past | pdf | discuss on HN

In plain words: A language model is trained to fill in missing words and name each word's meaning category from a dictionary, learning word senses without human labels. It chose word meanings more accurately than standard word-only training and beat previous systems on a word-in-context test.

Abstract · SenseBERT: Driving Some Sense into BERT

The ability to learn from large unlabeled corpora has allowed neural language models to advance the frontier in natural language understanding. However, existing self-supervision techniques operate at the word form level, which serves as a surrogate for the underlying semantic content. This paper proposes a method to employ weak-supervision directly at the word sense level. Our model, named SenseBERT, is pre-trained to predict not only the masked words but also their WordNet supersenses. Accordingly, we attain a lexical-semantic level language model, without the use of human annotation. SenseBERT achieves significantly improved lexical understanding, as we demonstrate by experimenting on SemEval Word Sense Disambiguation, and by attaining a state of the art result on the Word in Context task.

Yoav Levine, Barak Lenz, Or Dagan, Ori Ram, Dan Padnos, Or Sharir, Shai Shalev-Shwartz, Amnon Shashua, Yoav Shoham
arXiv:1908.05646 · cs.CL, cs.LG · submitted Aug 15, 2019 · updated May 18, 2020
abstract · pdf · html · Accepted to ACL 2020

add comment on HN