In plain words: It learns to translate using only separate text in each language, squeezing sentences from both into one shared space and rebuilding them in both. Unlike usual systems that need tens of thousands of matched sentence pairs, it scored 32.8 on a standard translation test.
Abstract · Unsupervised Machine Translation Using Monolingual Corpora Only
Machine translation has recently achieved impressive performance thanks to recent advances in deep learning and the availability of large-scale parallel corpora. There have been numerous attempts to extend these successes to low-resource language pairs, yet requiring tens of thousands of parallel sentences. In this work, we take this research direction to the extreme and investigate whether it is possible to learn to translate even without any parallel data. We propose a model that takes sentences from monolingual corpora in two different languages and maps them into the same latent space. By learning to reconstruct in both languages from this shared feature space, the model effectively learns to translate without using any labeled data. We demonstrate our model on two widely used datasets and two language pairs, reporting BLEU scores of 32.8 and 15.1 on the Multi30k and WMT English-French datasets, without using even a single parallel sentence at training time.
Guillaume Lample, Alexis Conneau, Ludovic Denoyer, Marc'Aurelio Ranzato
arXiv:1711.00043 · cs.CL, cs.AI · submitted Oct 31, 2017 · updated Apr 13, 2018
abstract · pdf · html · ICLR 2018
1- Train a system that translates language A sentences into a representation space, and can translate back from that space.
2- Train a second system that does the same, but with language B, onto the same representation space.
3- Train an adversarial system that tries to look at the representation space and identify which language the sentence came from, retraining the language translation systems to try to fool that recognizer. Retrain the models to try to not be recognized by this third system.
The best way to 'hide' from the recognizer is to have a very similar distribution, to make it use similar points in the representation space for the same concepts.
Brilliant stuff.