about
Distributed Representations of Sentences and Documents (arxiv.org)
5 points by sonabinu on Jan 10, 2019 | hide | past | pdf | 1 comment on HN

In plain words: Each document gets a fixed-length list of numbers, learned without labels by training it to predict the document's words. Unlike bag-of-words word counts, which ignore word order and meaning, it beat them and other text representations, achieving the best results on several classification and sentiment tasks.

Abstract

Many machine learning algorithms require the input to be represented as a fixed-length feature vector. When it comes to texts, one of the most common fixed-length features is bag-of-words. Despite their popularity, bag-of-words features have two major weaknesses: they lose the ordering of the words and they also ignore semantics of the words. For example, "powerful," "strong" and "Paris" are equally distant. In this paper, we propose Paragraph Vector, an unsupervised algorithm that learns fixed-length feature representations from variable-length pieces of texts, such as sentences, paragraphs, and documents. Our algorithm represents each document by a dense vector which is trained to predict words in the document. Its construction gives our algorithm the potential to overcome the weaknesses of bag-of-words models. Empirical results show that Paragraph Vectors outperform bag-of-words models as well as other techniques for text representations. Finally, we achieve new state-of-the-art results on several text classification and sentiment analysis tasks.

Quoc V. Le, Tomas Mikolov
arXiv:1405.4053 · cs.CL, cs.AI, cs.LG · submitted May 16, 2014 · updated May 22, 2014
abstract · pdf · html

add comment on HN
Also discussed: Dec 2016 (68 points, 13 comments) · May 2014 (1 point, 0 comments)

I like this idea much better than I like word vectors.

Although word vectors are well-established, at the end of the day people want to classify documents, not classify word vectors.

Also note that: (1) Sentiment analysis is where BoW goes to die (e.g. "not good" is roughly equal to "bad") and (2) I have not seen so many "beyond BoW" tasks other than sentiment analysis that have been well documented.