about
SMASH: One-Shot Model Architecture Search Through HyperNetworks (arxiv.org)
1 point by sirteno on Aug 29, 2017 | hide | past | pdf | discuss on HN

In plain words: A helper network learns to write out weights for any candidate design, so many architectures can be ranked after one training run instead of training each one separately. The best designs performed about as well as hand-built networks of similar size.

Abstract · SMASH: One-Shot Model Architecture Search through HyperNetworks

Designing architectures for deep neural networks requires expert knowledge and substantial computation time. We propose a technique to accelerate architecture selection by learning an auxiliary HyperNet that generates the weights of a main model conditioned on that model's architecture. By comparing the relative validation performance of networks with HyperNet-generated weights, we can effectively search over a wide range of architectures at the cost of a single training run. To facilitate this search, we develop a flexible mechanism based on memory read-writes that allows us to define a wide range of network connectivity patterns, with ResNet, DenseNet, and FractalNet blocks as special cases. We validate our method (SMASH) on CIFAR-10 and CIFAR-100, STL-10, ModelNet10, and Imagenet32x32, achieving competitive performance with similarly-sized hand-designed networks. Our code is available at https://github.com/ajbrock/SMASH

Andrew Brock, Theodore Lim, J. M. Ritchie, Nick Weston
arXiv:1708.05344 · cs.LG · submitted Aug 17, 2017
abstract · pdf · html

add comment on HN