about
Spectral Representations for Convolutional Neural Networks [pdf] (arxiv.org)
15 points by alexcasalboni on Jun 12, 2015 | hide | past | pdf | 2 comments on HN

In plain words: Instead of shrinking images by taking the biggest values in small blocks, this turns them into waves and keeps only the strongest, saving more detail per stored number. It also trains convolutional networks significantly faster while keeping competitive accuracy without max-pooling or dropout.

Abstract · Spectral Representations for Convolutional Neural Networks

Discrete Fourier transforms provide a significant speedup in the computation of convolutions in deep learning. In this work, we demonstrate that, beyond its advantages for efficient computation, the spectral domain also provides a powerful representation in which to model and train convolutional neural networks (CNNs). We employ spectral representations to introduce a number of innovations to CNN design. First, we propose spectral pooling, which performs dimensionality reduction by truncating the representation in the frequency domain. This approach preserves considerably more information per parameter than other pooling strategies and enables flexibility in the choice of pooling output dimensionality. This representation also enables a new form of stochastic regularization by randomized modification of resolution. We show that these methods achieve competitive results on classification and approximation tasks, without using any dropout or max-pooling. Finally, we demonstrate the effectiveness of complex-coefficient spectral parameterization of convolutional filters. While this leaves the underlying model unchanged, it results in a representation that greatly facilitates optimization. We observe on a variety of popular CNN configurations that this leads to significantly faster convergence during training.

Oren Rippel, Jasper Snoek, Ryan P. Adams
arXiv:1506.03767 · stat.ML, cs.LG · submitted Jun 11, 2015
abstract · pdf · html

add comment on HN

Reading through this paper, this is a logical extension of the recent work published by FAIR (Facebook AI Research) that proved that with careful implementation of FFT convolution that the speed up is very significant. The Facebook work though still had all the learning happening in the spatial not frequency domain. They are doing all the updating in the frequency domain, and have introduced a new type of pooling that uses stochastic resolution reduction in the frequency domain that seems very useful. Very interesting paper, I'm keen to try out the techniques myself.