In plain words: Wikipedia and web-crawl text were turned into word vectors—number lists that capture word meaning—for 157 languages, giving language apps a ready-made starting point. On the 10 languages with existing test sets, these vectors scored better than earlier ones.
Abstract
Distributed word representations, or word vectors, have recently been applied to many tasks in natural language processing, leading to state-of-the-art performance. A key ingredient to the successful application of these representations is to train them on very large corpora, and use these pre-trained models in downstream tasks. In this paper, we describe how we trained such high quality word representations for 157 languages. We used two sources of data to train these models: the free online encyclopedia Wikipedia and data from the common crawl project. We also introduce three new word analogy datasets to evaluate these word vectors, for French, Hindi and Polish. Finally, we evaluate our pre-trained word vectors on 10 languages for which evaluation datasets exists, showing very strong performance compared to previous models.
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, Tomas Mikolov
arXiv:1802.06893 · cs.CL, cs.LG · submitted Feb 19, 2018 · updated Mar 28, 2018
abstract · pdf · html · Accepted to LREC
Word vectors are fascinating representations. There is a huge amount of information and nuance captured in them. You can use them directly for topic retrieval (using annoy or another optimised vector index), or feed them into a classifier such as those in the Sklearn library. All types of neural nets: fully connected, recurrent and convolutional can be applied on word vectors.