about
Neural Machine Translation by Jointly Learning to Align and Translate (arxiv.org)
2 points by __Joker on Jan 25, 2015 | hide | past | pdf | discuss on HN

In plain words: Instead of squeezing the whole source sentence into one fixed vector, the model lets each translated word look back and pick out the relevant source words. On English-to-French it matched the best phrase-based system, and the word pairings it learned lined up with human intuition.

Abstract

Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance. The models proposed recently for neural machine translation often belong to a family of encoder-decoders and consists of an encoder that encodes a source sentence into a fixed-length vector from which a decoder generates a translation. In this paper, we conjecture that the use of a fixed-length vector is a bottleneck in improving the performance of this basic encoder-decoder architecture, and propose to extend this by allowing a model to automatically (soft-)search for parts of a source sentence that are relevant to predicting a target word, without having to form these parts as a hard segment explicitly. With this new approach, we achieve a translation performance comparable to the existing state-of-the-art phrase-based system on the task of English-to-French translation. Furthermore, qualitative analysis reveals that the (soft-)alignments found by the model agree well with our intuition.

Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio
arXiv:1409.0473 · cs.CL, cs.LG, cs.NE, stat.ML · submitted Sep 1, 2014 · updated May 19, 2016
abstract · pdf · html · Accepted at ICLR 2015 as oral presentation

add comment on HN
Also discussed: Apr 2023 (1 point, 0 comments)