about
Deep Networks with large output spaces – Faster training with hashes [pdf] (arxiv.org)
2 points by redknight666 on Dec 24, 2014 | hide | past | pdf | discuss on HN

In plain words: When a network must pick among millions of possible labels, scoring every one is too slow. A hashing trick quickly finds the few labels worth scoring by approximating the math behind each score, cutting training and prediction time compared with the usual check-everything approach.

Abstract · Deep Networks With Large Output Spaces

Deep neural networks have been extremely successful at various image, speech, video recognition tasks because of their ability to model deep structures within the data. However, they are still prohibitively expensive to train and apply for problems containing millions of classes in the output layer. Based on the observation that the key computation common to most neural network layers is a vector/matrix product, we propose a fast locality-sensitive hashing technique to approximate the actual dot product enabling us to scale up the training and inference to millions of output classes. We evaluate our technique on three diverse large-scale recognition tasks and show that our approach can train large-scale models at a faster rate (in terms of steps/total time) compared to baseline methods.

Sudheendra Vijayanarasimhan, Jonathon Shlens, Rajat Monga, Jay Yagnik
arXiv:1412.7479 · cs.NE, cs.LG · submitted Dec 23, 2014 · updated Apr 10, 2015
abstract · pdf · html

add comment on HN