about
Word2Vec Explained [pdf] (arxiv.org)
1 point by abishekk92 on Feb 2, 2015 | hide | past | pdf | discuss on HN

In plain words: This note unpacks the math behind word2vec's negative sampling, the trick that trains word embeddings by rewarding real word pairs and penalizing randomly picked fake ones. It walks through the derivation step by step, where the original papers left readers guessing.

Abstract · word2vec Explained: deriving Mikolov et al.'s negative-sampling word-embedding method

The word2vec software of Tomas Mikolov and colleagues (https://code.google.com/p/word2vec/ ) has gained a lot of traction lately, and provides state-of-the-art word embeddings. The learning models behind the software are described in two research papers. We found the description of the models in these papers to be somewhat cryptic and hard to follow. While the motivations and presentation may be obvious to the neural-networks language-modeling crowd, we had to struggle quite a bit to figure out the rationale behind the equations. This note is an attempt to explain equation (4) (negative sampling) in "Distributed Representations of Words and Phrases and their Compositionality" by Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado and Jeffrey Dean.

Yoav Goldberg, Omer Levy
arXiv:1402.3722 · cs.CL, cs.LG, stat.ML · submitted Feb 15, 2014
abstract · pdf · html

add comment on HN