about
Toward Geometric Deep SLAM (arxiv.org)
2 points by Impossible on Jul 26, 2017 | hide | past | pdf | discuss on HN

In plain words: Two small networks work together: one spots well-spread corner points in a photo, and the next compares two such point maps to figure out how the camera moved, using only point positions instead of the usual local descriptors. It beat classic point detectors in noisy images and ran 30+ frames per second on a CPU.

Abstract

We present a point tracking system powered by two deep convolutional neural networks. The first network, MagicPoint, operates on single images and extracts salient 2D points. The extracted points are "SLAM-ready" because they are by design isolated and well-distributed throughout the image. We compare this network against classical point detectors and discover a significant performance gap in the presence of image noise. As transformation estimation is more simple when the detected points are geometrically stable, we designed a second network, MagicWarp, which operates on pairs of point images (outputs of MagicPoint), and estimates the homography that relates the inputs. This transformation engine differs from traditional approaches because it does not use local point descriptors, only point locations. Both networks are trained with simple synthetic data, alleviating the requirement of expensive external camera ground truthing and advanced graphics rendering pipelines. The system is fast and lean, easily running 30+ FPS on a single CPU.

Daniel DeTone, Tomasz Malisiewicz, Andrew Rabinovich
arXiv:1707.07410 · cs.CV · submitted Jul 24, 2017
abstract · pdf · html

add comment on HN