In plain words: A single network scans an image and outputs keypoint locations and descriptors in one pass, teaching itself stable points by warping images instead of using hand labels. It finds far more repeatable points than corner detectors and aligns images better than SIFT or ORB.
Abstract · SuperPoint: Self-Supervised Interest Point Detection and Description
This paper presents a self-supervised framework for training interest point detectors and descriptors suitable for a large number of multiple-view geometry problems in computer vision. As opposed to patch-based neural networks, our fully-convolutional model operates on full-sized images and jointly computes pixel-level interest point locations and associated descriptors in one forward pass. We introduce Homographic Adaptation, a multi-scale, multi-homography approach for boosting interest point detection repeatability and performing cross-domain adaptation (e.g., synthetic-to-real). Our model, when trained on the MS-COCO generic image dataset using Homographic Adaptation, is able to repeatedly detect a much richer set of interest points than the initial pre-adapted deep model and any other traditional corner detector. The final system gives rise to state-of-the-art homography estimation results on HPatches when compared to LIFT, SIFT and ORB.
Daniel DeTone, Tomasz Malisiewicz, Andrew Rabinovich
arXiv:1712.07629 · cs.CV · submitted Dec 20, 2017 · updated Apr 19, 2018
abstract · pdf · html · Camera-ready version for CVPR 2018 Deep Learning for Visual SLAM Workshop (DL4VSLAM2018)