about
Rethinking Exposure Bias in Language Modeling (arxiv.org)
1 point by sel1 on Oct 27, 2019 | hide | past | pdf | discuss on HN

In plain words: Language models learn only from correct answers, so they stumble when fed their own mistakes. This one is trained on its own writing with rewards made stronger and less noisy, beating rivals on translation quality and a new test of handling its own errors.

Abstract · Rethinking Exposure Bias In Language Modeling

Exposure bias describes the phenomenon that a language model trained under the teacher forcing schema may perform poorly at the inference stage when its predictions are conditioned on its previous predictions unseen from the training corpus. Recently, several generative adversarial networks (GANs) and reinforcement learning (RL) methods have been introduced to alleviate this problem. Nonetheless, a common issue in RL and GANs training is the sparsity of reward signals. In this paper, we adopt two simple strategies, multi-range reinforcing, and multi-entropy sampling, to amplify and denoise the reward signal. Our model produces an improvement over competing models with regards to BLEU scores and road exam, a new metric we designed to measure the robustness against exposure bias in language models.

Yifan Xu, Kening Zhang, Haoyu Dong, Yuezhou Sun, Wenlong Zhao, Zhuowen Tu
arXiv:1910.11235 · cs.CL, cs.LG · submitted Oct 13, 2019 · updated Mar 31, 2020
abstract · pdf · html

add comment on HN