In plain words: Measuring accuracy, memory, size, operation count, speed and power use across top image-recognition models shows what each really costs. Power draw stayed the same no matter the batch size or architecture, while higher accuracy demanded sharply longer waits.
Abstract
Since the emergence of Deep Neural Networks (DNNs) as a prominent technique in the field of computer vision, the ImageNet classification challenge has played a major role in advancing the state-of-the-art. While accuracy figures have steadily increased, the resource utilisation of winning models has not been properly taken into account. In this work, we present a comprehensive analysis of important metrics in practical applications: accuracy, memory footprint, parameters, operations count, inference time and power consumption. Key findings are: (1) power consumption is independent of batch size and architecture; (2) accuracy and inference time are in a hyperbolic relationship; (3) energy constraint is an upper bound on the maximum achievable accuracy and model complexity; (4) the number of operations is a reliable estimate of the inference time. We believe our analysis provides a compelling set of information that helps design and engineer efficient DNNs.
Alfredo Canziani, Adam Paszke, Eugenio Culurciello
arXiv:1605.07678 · cs.CV · submitted May 24, 2016 · updated Apr 14, 2017
abstract · pdf · html · 7 pages, 10 figures, legend for Figure 2 got lost :/