about
Span Selection Pre-Training for Question Answering (arxiv.org)
1 point by sel1 on Sep 12, 2019 | hide | past | pdf | discuss on HN

In plain words: Instead of pre-training by guessing missing words from memory, it learns to pick the answer out of the passage, like a reading-comprehension question. It beat the usual large pre-trained model by 3 points on short answers and helped most when training data was scarce.

Abstract · Span Selection Pre-training for Question Answering

BERT (Bidirectional Encoder Representations from Transformers) and related pre-trained Transformers have provided large gains across many language understanding tasks, achieving a new state-of-the-art (SOTA). BERT is pre-trained on two auxiliary tasks: Masked Language Model and Next Sentence Prediction. In this paper we introduce a new pre-training task inspired by reading comprehension to better align the pre-training from memorization to understanding. Span Selection Pre-Training (SSPT) poses cloze-like training instances, but rather than draw the answer from the model's parameters, it is selected from a relevant passage. We find significant and consistent improvements over both BERT-BASE and BERT-LARGE on multiple reading comprehension (MRC) datasets. Specifically, our proposed model has strong empirical evidence as it obtains SOTA results on Natural Questions, a new benchmark MRC dataset, outperforming BERT-LARGE by 3 F1 points on short answer prediction. We also show significant impact in HotpotQA, improving answer prediction F1 by 4 points and supporting fact prediction F1 by 1 point and outperforming the previous best system. Moreover, we show that our pre-training approach is particularly effective when training data is limited, improving the learning curve by a large amount.

Michael Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto, Lin Pan, G P Shrivatsa Bhargav, Dinesh Garg, Avirup Sil
arXiv:1909.04120 · cs.CL, cs.AI, cs.LG · submitted Sep 9, 2019 · updated Jun 18, 2020
abstract · pdf · html · Accepted at ACL2020

add comment on HN