In plain words: A reference that works out how a layer's input size, filter size, padding, and step size set its output size, for regular, pooling, and upsampling layers. Instead of guessing and testing, you can plug numbers into the derived formulas to size a network correctly.
Abstract
We introduce a guide to help deep learning practitioners understand and manipulate convolutional neural network architectures. The guide clarifies the relationship between various properties (input shape, kernel shape, zero padding, strides and output shape) of convolutional, pooling and transposed convolutional layers, as well as the relationship between convolutional and transposed convolutional layers. Relationships are derived for various cases, and are illustrated in order to make them intuitive.
Vincent Dumoulin, Francesco Visin
arXiv:1603.07285 · stat.ML, cs.LG, cs.NE · submitted Mar 23, 2016 · updated Jan 11, 2018
abstract · pdf · html