about
Semantics-Aware Bert for Language Understanding (arxiv.org)
2 points by sel1 on Sep 8, 2019 | hide | past | pdf | 1 comment on HN

In plain words: This system layers extra meaning clues—who did what to whom, marked by a separate tool—onto a standard word-context model to help it understand language better. It beat the plain version it builds on across ten reading and reasoning tasks.

Abstract · Semantics-aware BERT for Language Understanding

The latest work on language representations carefully integrates contextualized features into language model training, which enables a series of success especially in various machine reading comprehension and natural language inference tasks. However, the existing language representation models including ELMo, GPT and BERT only exploit plain context-sensitive features such as character or word embeddings. They rarely consider incorporating structured semantic information which can provide rich semantics for language representation. To promote natural language understanding, we propose to incorporate explicit contextual semantics from pre-trained semantic role labeling, and introduce an improved language representation model, Semantics-aware BERT (SemBERT), which is capable of explicitly absorbing contextual semantics over a BERT backbone. SemBERT keeps the convenient usability of its BERT precursor in a light fine-tuning way without substantial task-specific modifications. Compared with BERT, semantics-aware BERT is as simple in concept but more powerful. It obtains new state-of-the-art or substantially improves results on ten reading comprehension and language inference tasks.

Zhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li, Shuailiang Zhang, Xi Zhou, Xiang Zhou
arXiv:1909.02209 · cs.CL · submitted Sep 5, 2019 · updated Feb 4, 2020
abstract · pdf · html · Thirty-Fourth AAAI Conference on Artificial Intelligence (AAAI-2020)

add comment on HN

Nice paper! I will never forget my surprise earlier this year about how well BERT handled coreference (anaphoric resolution) tasks. Adding structured context data sounds good, but I need to figure out from the paper how it works.