In plain words: The model learns sentence meanings straight from pairs of the same text in two languages, skipping the usual step of hand-building lists of reworded sentences. On tasks comparing text across languages, it beat more complicated systems while running orders of magnitude faster.
Abstract
We present a model and methodology for learning paraphrastic sentence embeddings directly from bitext, removing the time-consuming intermediate step of creating paraphrase corpora. Further, we show that the resulting model can be applied to cross-lingual tasks where it both outperforms and is orders of magnitude faster than more complex state-of-the-art baselines.
John Wieting, Kevin Gimpel, Graham Neubig, Taylor Berg-Kirkpatrick
arXiv:1909.13872 · cs.CL · submitted Sep 30, 2019
abstract · pdf · html · Published as a short paper at ACL 2019