In plain words: A translation model that reads source words and writes target words in one stream, starting output after the first input word and using constant memory. It matched the usual model where output words look back at the input, and did better on long sentences.
Abstract · You May Not Need Attention
In NMT, how far can we get without attention and without separate encoding and decoding? To answer that question, we introduce a recurrent neural translation model that does not use attention and does not have a separate encoder and decoder. Our eager translation model is low-latency, writing target tokens as soon as it reads the first source token, and uses constant memory during decoding. It performs on par with the standard attention-based model of Bahdanau et al. (2014), and better on long sentences.
Ofir Press, Noah A. Smith
arXiv:1810.13409 · cs.CL · submitted Oct 31, 2018
abstract · pdf · html