In plain words: Words are weighted by how far their vectors sit from the spread of nearby words, so odd words count more; weighted word counts form a sentence vector. It beat rarity weighting on every test and a trained encoder on most, with almost no training.
Abstract · Contextual Salience for Fast and Accurate Sentence Vectors
Unsupervised vector representations of sentences or documents are a major building block for many language tasks such as sentiment classification. However, current methods are uninterpretable and slow or require large training datasets. Recent word vector-based proposals implicitly assume that distances in a word embedding space are equally important, regardless of context. We introduce contextual salience (CoSal), a measure of word importance that uses the distribution of context vectors to normalize distances and weights. CoSal relies on the insight that unusual word vectors disproportionately affect phrase vectors. A bag-of-words model with CoSal-based weights produces accurate unsupervised sentence or document representations for classification, requiring little computation to evaluate and only a single covariance calculation to ``train." CoSal supports small contexts, out-of context words and outperforms SkipThought on most benchmarks, beats tf-idf on all benchmarks, and is competitive with the unsupervised state-of-the-art.
Eric Zelikman, Richard Socher
arXiv:1803.08493 · cs.CL · submitted Mar 22, 2018 · updated Nov 2, 2020
abstract · pdf · html
TL;DR:
1. Build a special distance(word1, word2) metric using a set of word vectors trained elsewhere (such as GloVe). This distance works better than cosine distance.
2. Given a document (a sentence, a paragraph… basically, a sequence of words), calculate the "importance" of each word as a sigmoid over distances(word, avg(all_words)).
3. To embed a document, simply do a weighted average of its word vectors, where the weight of each word equals the importance above.