about
Group Normalization (arxiv.org)
82 points by victorvation on Mar 24, 2018 | hide | past | pdf | 12 comments on HN

In plain words: Instead of averaging across a batch of images, this normalizes each image's channels in small groups, so accuracy doesn't depend on batch size. With just 2 images per batch it cut image-recognition error 10.6% versus batch normalization, and matched it at normal batch sizes.

Abstract

Batch Normalization (BN) is a milestone technique in the development of deep learning, enabling various networks to train. However, normalizing along the batch dimension introduces problems --- BN's error increases rapidly when the batch size becomes smaller, caused by inaccurate batch statistics estimation. This limits BN's usage for training larger models and transferring features to computer vision tasks including detection, segmentation, and video, which require small batches constrained by memory consumption. In this paper, we present Group Normalization (GN) as a simple alternative to BN. GN divides the channels into groups and computes within each group the mean and variance for normalization. GN's computation is independent of batch sizes, and its accuracy is stable in a wide range of batch sizes. On ResNet-50 trained in ImageNet, GN has 10.6% lower error than its BN counterpart when using a batch size of 2; when using typical batch sizes, GN is comparably good with BN and outperforms other normalization variants. Moreover, GN can be naturally transferred from pre-training to fine-tuning. GN can outperform its BN-based counterparts for object detection and segmentation in COCO, and for video classification in Kinetics, showing that GN can effectively replace the powerful BN in a variety of tasks. GN can be easily implemented by a few lines of code in modern libraries.

Yuxin Wu, Kaiming He
arXiv:1803.08494 · cs.CV, cs.LG · submitted Mar 22, 2018 · updated Jun 11, 2018
abstract · pdf · html · v3: Update trained-from-scratch results in COCO to 41.0AP. Code and models at https://github.com/facebookresearch/Detectron/blob/master/projects/GN

add comment on HN

No relation of subgroup normalisation, just a bad name.
Mostly only relevant for convnets it seems
"only"
Has anyone made or found an implementation yet?
When it comes to using super new research models or methods in deep learning, implementing it yourself instead of looking for other implementations is almost a requirement. In my experience the majority of online "implementations" of research deep learning models or methods have subtle bugs, flaws, or outright implemented something different than the paper they reference. This isn't limited to random github code: for example, I wouldn't trust anything in tensorflow.contrib unless I read the source.
Having implemented algorithms from research papers I came to the conclusion that the same is true about the papers. It’s unbelievable how bad the quality of the description of algorithms often is.
Are you generally able to achieve the results the new research claims after you implement it yourself per their description?

I would have thought if their description is shit then their results are shit too (cherry-picked from hundreds of runs, outright fabricated, misinterpreted, incorrectly graphed, whatever - just plain doesn't work, like a bad recipe.)

As a followup if this is the case then how do you know what new deep learning research is worth your time to implement?

Completely agree. Well, tf.contrib I am usually happy with the docs and a glance through GitHub issues, but, yes, code found elsewhere for very new, or even just recent papers can be really hit or miss, and not always obviously so. There is a lot of great stuff out there, but sometimes even popular papers over a year old can still be a challenge to find good code for. finding 3 separate implementations, all stlighty wrong or unusable is not uncommon.
available in PyTorch master via: https://github.com/pytorch/pytorch/pull/5968
Figure 3 in the paper.
Nothing to do with database normalization or reconciliation