about
Evolving Normalization-Activation Layers (arxiv.org)
2 points by memexy on Jun 14, 2020 | hide | past | pdf | 1 comment on HN

In plain words: A computer search builds one layer that does both normalization and activation, combining math steps like addition and multiplication instead of stacking separate pieces. The discovered layers beat the batch and group normalization setups on image tasks and transfer to detection and image generation.

Abstract

Normalization layers and activation functions are fundamental components in deep networks and typically co-locate with each other. Here we propose to design them using an automated approach. Instead of designing them separately, we unify them into a single tensor-to-tensor computation graph, and evolve its structure starting from basic mathematical functions. Examples of such mathematical functions are addition, multiplication and statistical moments. The use of low-level mathematical functions, in contrast to the use of high-level modules in mainstream NAS, leads to a highly sparse and large search space which can be challenging for search methods. To address the challenge, we develop efficient rejection protocols to quickly filter out candidate layers that do not work well. We also use multi-objective evolution to optimize each layer's performance across many architectures to prevent overfitting. Our method leads to the discovery of EvoNorms, a set of new normalization-activation layers with novel, and sometimes surprising structures that go beyond existing design patterns. For example, some EvoNorms do not assume that normalization and activation functions must be applied sequentially, nor need to center the feature maps, nor require explicit activation functions. Our experiments show that EvoNorms work well on image classification models including ResNets, MobileNets and EfficientNets but also transfer well to Mask R-CNN with FPN/SpineNet for instance segmentation and to BigGAN for image synthesis, outperforming BatchNorm and GroupNorm based layers in many cases.

Hanxiao Liu, Andrew Brock, Karen Simonyan, Quoc V. Le
arXiv:2004.02967 · cs.LG, cs.CV, cs.NE, stat.ML · submitted Apr 6, 2020 · updated Jul 17, 2020
abstract · pdf · html

add comment on HN
Also discussed: Apr 2020 (2 points, 0 comments)

> Normalization layers and activation functions are critical components in deep neural networks that frequently co-locate with each other. Instead of designing them separately, we unify them into a single computation graph, and evolve its structure starting from low-level primitives. Our layer search algorithm leads to the discovery of EvoNorms, a set of new normalization-activation layers that go beyond existing design patterns. Several of these layers enjoy the property of being independent from the batch statistics. Our experiments show that EvoNorms not only work well on a variety of image classification models including ResNets, MobileNets and EfficientNets but also transfer well to Mask R-CNN, SpineNet for instance segmentation and BigGAN for image synthesis, significantly outperforming BatchNorm and GroupNorm based layers in many cases.

I recently was wondering why evolutionary tactics are not used to evolve neural network architectures and this paper answers that question. People are looking at evolving novel architectures using evolutionary tactics.