about
Breaking the Softmax Bottleneck: A High-Rank RNN Language Model (arxiv.org)
4 points by stablemap on Nov 13, 2017 | hide | past | pdf | discuss on HN

In plain words: Standard next-word predictors use one fixed set of word vectors, which limits how many context-dependent word patterns they can express. Mixing several prediction layers lifts that limit, cutting the error score (lower is better) to 47.69 on a text dataset, beating the previous best.

Abstract

We formulate language modeling as a matrix factorization problem, and show that the expressiveness of Softmax-based models (including the majority of neural language models) is limited by a Softmax bottleneck. Given that natural language is highly context-dependent, this further implies that in practice Softmax with distributed word embeddings does not have enough capacity to model natural language. We propose a simple and effective method to address this issue, and improve the state-of-the-art perplexities on Penn Treebank and WikiText-2 to 47.69 and 40.68 respectively. The proposed method also excels on the large-scale 1B Word dataset, outperforming the baseline by over 5.6 points in perplexity.

Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, William W. Cohen
arXiv:1711.03953 · cs.CL, cs.LG · submitted Nov 10, 2017 · updated Mar 2, 2018
abstract · pdf · html · ICLR Oral 2018

add comment on HN