about
A guide to convolution arithmetic for deep learning (arxiv.org)
2 points by bjourne on Aug 12, 2022 | hide | past | pdf | discuss on HN

In plain words: A reference that works out how a layer's input size, filter size, padding, and step size set its output size, for regular, pooling, and upsampling layers. Instead of guessing and testing, you can plug numbers into the derived formulas to size a network correctly.

Abstract

We introduce a guide to help deep learning practitioners understand and manipulate convolutional neural network architectures. The guide clarifies the relationship between various properties (input shape, kernel shape, zero padding, strides and output shape) of convolutional, pooling and transposed convolutional layers, as well as the relationship between convolutional and transposed convolutional layers. Relationships are derived for various cases, and are illustrated in order to make them intuitive.

Vincent Dumoulin, Francesco Visin
arXiv:1603.07285 · stat.ML, cs.LG, cs.NE · submitted Mar 23, 2016 · updated Jan 11, 2018
abstract · pdf · html

add comment on HN
Also discussed: Jul 2020 (1 point, 0 comments) · Feb 2020 (4 points, 0 comments) · Jan 2018 (2 points, 0 comments) · May 2017 (3 points, 0 comments) · Mar 2016 (6 points, 0 comments)