In plain words: A word-by-word language model gains an attention step that scores phrases in earlier text, and those scores can be read as a grammar tree of the sentence. It predicted text better than strong rivals, and teaching the attention step with grammar labels helped more.
Abstract · PaLM: A Hybrid Parser and Language Model
We present PaLM, a hybrid parser and neural language model. Building on an RNN language model, PaLM adds an attention layer over text spans in the left context. An unsupervised constituency parser can be derived from its attention weights, using a greedy decoding algorithm. We evaluate PaLM on language modeling, and empirically show that it outperforms strong baselines. If syntactic annotations are available, the attention component can be trained in a supervised manner, providing syntactically-informed representations of the context, and further improving language modeling performance.
Hao Peng, Roy Schwartz, Noah A. Smith
arXiv:1909.02134 · cs.CL · submitted Sep 4, 2019
abstract · pdf · html · EMNLP 2019 short paper