In plain words: A single neural network finds interesting points in photos, works out how they are turned, and describes them, with all three steps trained together so the whole pipeline learns from mistakes. It beat the best hand-designed methods on several image benchmarks without retraining.
Abstract
We introduce a novel Deep Network architecture that implements the full feature point handling pipeline, that is, detection, orientation estimation, and feature description. While previous works have successfully tackled each one of these problems individually, we show how to learn to do all three in a unified manner while preserving end-to-end differentiability. We then demonstrate that our Deep pipeline outperforms state-of-the-art methods on a number of benchmark datasets, without the need of retraining.
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, Pascal Fua
arXiv:1603.09114 · cs.CV · submitted Mar 30, 2016 · updated Jul 29, 2016
abstract · pdf · html · Accepted to ECCV 2016 (spotlight)