about
Parameter Prediction for Unseen Deep Architectures (arxiv.org)
2 points by jedwhite on Mar 10, 2022 | hide | past | pdf | discuss on HN

In plain words: Instead of slowly tuning a network's weights, this system reads its wiring diagram and spits out all the weights in one quick pass. It guessed all 24 million weights of a ResNet-50 well enough to hit 60% accuracy on CIFAR-10, even on unseen networks.

Abstract

Deep learning has been successful in automating the design of features in machine learning pipelines. However, the algorithms optimizing neural network parameters remain largely hand-designed and computationally inefficient. We study if we can use deep learning to directly predict these parameters by exploiting the past knowledge of training other networks. We introduce a large-scale dataset of diverse computational graphs of neural architectures - DeepNets-1M - and use it to explore parameter prediction on CIFAR-10 and ImageNet. By leveraging advances in graph neural networks, we propose a hypernetwork that can predict performant parameters in a single forward pass taking a fraction of a second, even on a CPU. The proposed model achieves surprisingly good performance on unseen and diverse networks. For example, it is able to predict all 24 million parameters of a ResNet-50 achieving a 60% accuracy on CIFAR-10. On ImageNet, top-5 accuracy of some of our networks approaches 50%. Our task along with the model and results can potentially lead to a new, more computationally efficient paradigm of training networks. Our model also learns a strong representation of neural architectures enabling their analysis.

Boris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-Soriano
arXiv:2110.13100 · cs.LG, cs.AI, cs.CV, stat.ML · submitted Oct 25, 2021
abstract · pdf · html · NeurIPS 2021 camera ready, the code is available at https://github.com/facebookresearch/ppuda

add comment on HN
Also discussed: Nov 2021 (1 point, 1 comment)