about
Molding CNNs for text: non-linear, non-consecutive convolutions (arxiv.org)
16 points by titocosta on Aug 18, 2015 | hide | past | pdf | discuss on HN

In plain words: Instead of sliding a window over adjacent words and gluing vectors together, this text classifier multiplies word vectors to capture interactions and can skip words in between. It beat the usual approach on sentiment and news tasks, reaching 51.2% on fine-grained sentiment and training faster.

Abstract

The success of deep learning often derives from well-chosen operational building blocks. In this work, we revise the temporal convolution operation in CNNs to better adapt it to text processing. Instead of concatenating word representations, we appeal to tensor algebra and use low-rank n-gram tensors to directly exploit interactions between words already at the convolution stage. Moreover, we extend the n-gram convolution to non-consecutive words to recognize patterns with intervening words. Through a combination of low-rank tensors, and pattern weighting, we can efficiently evaluate the resulting convolution operation via dynamic programming. We test the resulting architecture on standard sentiment classification and news categorization tasks. Our model achieves state-of-the-art performance both in terms of accuracy and training speed. For instance, we obtain 51.2% accuracy on the fine-grained sentiment classification task.

Tao Lei, Regina Barzilay, Tommi Jaakkola
arXiv:1508.04112 · cs.CL, cs.AI · submitted Aug 17, 2015 · updated Aug 18, 2015
abstract · pdf · html

add comment on HN