In plain words: Word vectors are lists of numbers that capture what words mean, learned from lots of text. Combining several known training tricks that are rarely used together produced new public word vectors that beat the best current ones by a large margin on many tasks.
Abstract
Many Natural Language Processing applications nowadays rely on pre-trained word representations estimated from large text corpora such as news collections, Wikipedia and Web Crawl. In this paper, we show how to train high-quality word vector representations by using a combination of known tricks that are however rarely used together. The main result of our work is the new set of publicly available pre-trained models that outperform the current state of the art by a large margin on a number of tasks.
Tomas Mikolov, Edouard Grave, Piotr Bojanowski, Christian Puhrsch, Armand Joulin
arXiv:1712.09405 · cs.CL · submitted Dec 26, 2017
abstract · pdf · html
As they write in the abstract, they don't introduce any new techniques, but I think the paper reads really well as a concise tour for people interested in the current state of the art.