about
ConvNet Features off-the-shelf: an Astounding Baseline for Object Recognition (arxiv.org)
10 points by m_ke on Apr 4, 2014 | hide | past | pdf | 4 comments on HN

In plain words: Take the features of a network trained only to sort objects into categories, then use a simple classifier to recognize scenes, details, or attributes. They beat highly tuned task-specific systems, and for image search they beat other memory-light methods except on sculptures.

Abstract · CNN Features off-the-shelf: an Astounding Baseline for Recognition

Recent results indicate that the generic descriptors extracted from the convolutional neural networks are very powerful. This paper adds to the mounting evidence that this is indeed the case. We report on a series of experiments conducted for different recognition tasks using the publicly available code and model of the \overfeat network which was trained to perform object classification on ILSVRC13. We use features extracted from the \overfeat network as a generic image representation to tackle the diverse range of recognition tasks of object image classification, scene recognition, fine grained recognition, attribute detection and image retrieval applied to a diverse set of datasets. We selected these tasks and datasets as they gradually move further away from the original task and data the \overfeat network was trained to solve. Astonishingly, we report consistent superior results compared to the highly tuned state-of-the-art systems in all the visual classification tasks on various datasets. For instance retrieval it consistently outperforms low memory footprint methods except for sculptures dataset. The results are achieved using a linear SVM classifier (or $L2$ distance in case of retrieval) applied to a feature representation of size 4096 extracted from a layer in the net. The representations are further modified using simple augmentation techniques e.g. jittering. The results strongly suggest that features obtained from deep learning with convolutional nets should be the primary candidate in most visual recognition tasks.

Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, Stefan Carlsson
arXiv:1403.6382 · cs.CV · submitted Mar 23, 2014 · updated May 12, 2014
abstract · pdf · html · version 3 revisions: 1)Added results using feature processing and data augmentation 2)Referring to most recent efforts of using CNN for different visual recognition tasks 3) updated text/caption

add comment on HN

Quick summary for people who don't follow machine learning and computer vision:

A team from NYU recently open sourced their neural net model from a recent object recognition competition. The team from KTH used that pre-trained net as a feature extractor and applied it to other standard datasets by learning an SVM on the extracted features. This "simple" approach did as well if not better than most state of the art methods, which were hand engineered for the specific tasks.

CNNs are great, but so prone to overfitting, thankfully that last few years have been good to the deep learning community - dropout, jitter, etc. - in combating these problems. Hopefully in the next few years more books are published to help dumb-down the material since reading a lot of these papers (in terms of trying to implement/expand/critique the material - just the history alone is interesting - hopfield nets, perceptrons, hebbian learning, ... ) is still very tough as results are difficult to reproduce
There is a nice python wrapper, called nolearn [1] that uses a pre-trained CNN called DeCaF to extract features prior to the final classification layer. I found the results to be surprisingly good in a lot of situations.

[1]: https://pythonhosted.org/nolearn/

I have to wonder - are there some techniques to build the trained models by crowdsourcing?