about
SDNet: Semantically Guided Depth Estimation Network (arxiv.org)
3 points by sel1 on Jul 28, 2019 | hide | past | pdf | discuss on HN

In plain words: A single network predicts each pixel's depth and what it is at the same time, instead of running two separate networks, and guesses depth by ranking it into ordered distance bands. It beat the two-network setup on both tasks while using less computing power.

Abstract

Autonomous vehicles and robots require a full scene understanding of the environment to interact with it. Such a perception typically incorporates pixel-wise knowledge of the depths and semantic labels for each image from a video sensor. Recent learning-based methods estimate both types of information independently using two separate CNNs. In this paper, we propose a model that is able to predict both outputs simultaneously, which leads to improved results and even reduced computational costs compared to independent estimation of depth and semantics. We also empirically prove that the CNN is capable of learning more meaningful and semantically richer features. Furthermore, our SDNet estimates the depth based on ordinal classification. On the basis of these two enhancements, our proposed method achieves state-of-the-art results in semantic segmentation and depth estimation from single monocular input images on two challenging datasets.

Matthias Ochs, Adrian Kretz, Rudolf Mester
arXiv:1907.10659 · cs.CV · submitted Jul 24, 2019
abstract · pdf · html · Paper is accepted at German Conference on Pattern Recognition (GCPR), Dortmund, Germany, September 2019

add comment on HN