In plain words: Each word gets vectors, one per meaning like noun or verb, and a model picks the right one using supervised disambiguation, cheaper to train than clustering words by context. In dependency parsing, it cut unlabeled attachment errors by over 8% averaged across 6 languages.
Abstract · sense2vec - A Fast and Accurate Method for Word Sense Disambiguation In Neural Word Embeddings
Neural word representations have proven useful in Natural Language Processing (NLP) tasks due to their ability to efficiently model complex semantic and syntactic word relationships. However, most techniques model only one representation per word, despite the fact that a single word can have multiple meanings or "senses". Some techniques model words by using multiple vectors that are clustered based on context. However, recent neural approaches rarely focus on the application to a consuming NLP algorithm. Furthermore, the training process of recent word-sense models is expensive relative to single-sense embedding processes. This paper presents a novel approach which addresses these concerns by modeling multiple embeddings for each word based on supervised disambiguation, which provides a fast and accurate way for a consuming NLP model to select a sense-disambiguated embedding. We demonstrate that these embeddings can disambiguate both contrastive senses such as nominal and verbal senses as well as nuanced senses such as sarcasm. We further evaluate Part-of-Speech disambiguated embeddings on neural dependency parsing, yielding a greater than 8% average error reduction in unlabeled attachment scores across 6 languages.
Andrew Trask, Phil Michalak, John Liu
arXiv:1511.06388 · cs.CL, cs.LG · submitted Nov 19, 2015
abstract · pdf · html
To understand the technique, first understand word2vec:
http://rare-technologies.com/word2vec-tutorial/
http://colah.github.io/posts/2014-07-NLP-RNNs-Representation...
Now understand part-of-speech tagging:
http://spacy.io/blog/part-of-speech-POS-tagger-in-python/
By default word2vec gives you clusters for each word, this paper is giving you clusters for word_POS, e.g. The_DT apple_NNP employee_NN is_VBZ eating_VBG an_DT apple_NN. The same trick is done with named entity labels as well.
The following papers explain how the new word vectors are used in a dependency parser:
Collobert and Weston (2011): http://arxiv.org/pdf/1103.0398.pdf
Wang and Manning (2014): http://cs.stanford.edu/~danqi/papers/emnlp2014.pdf
Yoav Goldberg (2015): http://u.cs.biu.ac.il/~yogo/nnlp.pdf Survey/review, aimed at grad students