about
Grammar as a Foreign Language (Google Research) (arxiv.org)
1 point by bra-ket on Sep 19, 2015 | hide | past | pdf | discuss on HN

In plain words: It turns a sentence into a tree written as brackets and labels, using a translation-style network that lets the output look back at the input. Trained on machine-annotated examples, it beat specialized parsers on the standard test set and ran over 100 sentences a second.

Abstract · Grammar as a Foreign Language

Syntactic constituency parsing is a fundamental problem in natural language processing and has been the subject of intensive research and engineering for decades. As a result, the most accurate parsers are domain specific, complex, and inefficient. In this paper we show that the domain agnostic attention-enhanced sequence-to-sequence model achieves state-of-the-art results on the most widely used syntactic constituency parsing dataset, when trained on a large synthetic corpus that was annotated using existing parsers. It also matches the performance of standard parsers when trained only on a small human-annotated dataset, which shows that this model is highly data-efficient, in contrast to sequence-to-sequence models without the attention mechanism. Our parser is also fast, processing over a hundred sentences per second with an unoptimized CPU implementation.

Oriol Vinyals, Lukasz Kaiser, Terry Koo, Slav Petrov, Ilya Sutskever, Geoffrey Hinton
arXiv:1412.7449 · cs.CL, cs.LG, stat.ML · submitted Dec 23, 2014 · updated Jun 9, 2015
abstract · pdf · html

add comment on HN