about
The Devil is in the Decoder (arxiv.org)
2 points by lainon on Jul 20, 2017 | hide | past | pdf | discuss on HN

In plain words: Vision models shrink an image to learn features, then use a decoder to rebuild a full-size prediction for every pixel. Testing many decoder designs across segmentation, depth prediction, and image generation showed the choice changes results a lot; new upsampling and skip connections were also introduced.

Abstract · The Devil is in the Decoder: Classification, Regression and GANs

Many machine vision applications, such as semantic segmentation and depth prediction, require predictions for every pixel of the input image. Models for such problems usually consist of encoders which decrease spatial resolution while learning a high-dimensional representation, followed by decoders who recover the original input resolution and result in low-dimensional predictions. While encoders have been studied rigorously, relatively few studies address the decoder side. This paper presents an extensive comparison of a variety of decoders for a variety of pixel-wise tasks ranging from classification, regression to synthesis. Our contributions are: (1) Decoders matter: we observe significant variance in results between different types of decoders on various problems. (2) We introduce new residual-like connections for decoders. (3) We introduce a novel decoder: bilinear additive upsampling. (4) We explore prediction artifacts.

Zbigniew Wojna, Vittorio Ferrari, Sergio Guadarrama, Nathan Silberman, Liang-Chieh Chen, Alireza Fathi, Jasper Uijlings
arXiv:1707.05847 · cs.CV · submitted Jul 18, 2017 · updated Feb 19, 2019
abstract · pdf · html

add comment on HN