about
FFT Convolutions Are Faster Than Winograd on Modern CPUs, Here Is Why (arxiv.org)
3 points by Katydid on Sep 30, 2018 | hide | past | pdf | discuss on HN

In plain words: They timed three ways to run convolutions on modern CPUs—two using the Fourier transform, one that cuts arithmetic—and checked memory traffic and cache use, not just operation counts. The Fourier versions were faster, showing memory and cache limits matter more than math.

Abstract · FFT Convolutions are Faster than Winograd on Modern CPUs, Here is Why

Winograd-based convolution has quickly gained traction as a preferred approach to implement convolutional neural networks (ConvNet) on various hardware platforms because it requires fewer floating point operations than FFT-based or direct convolutions. This paper compares three highly optimized implementations (regular FFT--, Gauss--FFT--, and Winograd--based convolutions) on modern multi-- and many--core CPUs. Although all three implementations employed the same optimizations for modern CPUs, our experimental results with two popular ConvNets (VGG and AlexNet) show that the FFT--based implementations generally outperform the Winograd--based approach, contrary to the popular belief. To understand the results, we use a Roofline performance model to analyze the three implementations in detail, by looking at each of their computation phases and by considering not only the number of floating point operations, but also the memory bandwidth and the cache sizes. The performance analysis explains why, and under what conditions, the FFT--based implementations outperform the Winograd--based one, on modern CPUs.

Aleksandar Zlateski, Zhen Jia, Kai Li, Fredo Durand
arXiv:1809.07851 · cs.PF · submitted Sep 20, 2018
abstract · pdf · html

add comment on HN