about
VanillaNet: The power of minimalism in deep learning (arxiv.org)
33 points by pizza on May 24, 2023 | hide | past | pdf | 5 comments on HN

In plain words: A very shallow network built from simple layers, with no skip connections or attention; it trains with extra nonlinear steps, then removes them so the final design stays plain and fast. It matched the accuracy of deep networks and vision transformers while remaining far simpler.

Abstract · VanillaNet: the Power of Minimalism in Deep Learning

At the heart of foundation models is the philosophy of "more is different", exemplified by the astonishing success in computer vision and natural language processing. However, the challenges of optimization and inherent complexity of transformer models call for a paradigm shift towards simplicity. In this study, we introduce VanillaNet, a neural network architecture that embraces elegance in design. By avoiding high depth, shortcuts, and intricate operations like self-attention, VanillaNet is refreshingly concise yet remarkably powerful. Each layer is carefully crafted to be compact and straightforward, with nonlinear activation functions pruned after training to restore the original architecture. VanillaNet overcomes the challenges of inherent complexity, making it ideal for resource-constrained environments. Its easy-to-understand and highly simplified architecture opens new possibilities for efficient deployment. Extensive experimentation demonstrates that VanillaNet delivers performance on par with renowned deep neural networks and vision transformers, showcasing the power of minimalism in deep learning. This visionary journey of VanillaNet has significant potential to redefine the landscape and challenge the status quo of foundation model, setting a new path for elegant and effective model design. Pre-trained models and codes are available at https://github.com/huawei-noah/VanillaNet and https://gitee.com/mindspore/models/tree/master/research/cv/vanillanet.

Hanting Chen, Yunhe Wang, Jianyuan Guo, Dacheng Tao
arXiv:2305.12972 · cs.CV · submitted May 22, 2023 · updated May 23, 2023
abstract · pdf · html

add comment on HN

I too am repelled by the complexity of deep learning research, but using the phrase "visionary journey" in the abstract is a bit of a crank alert...
The whole abstract sounds like marketing copy from those websites that don't tell you what their product is:

> By avoiding high depth, shortcuts, and intricate operations like self-attention, VanillaNet is refreshingly concise yet remarkably powerful.

This seems bad, though?

They introduce a complicated training regime to make up for tearing out a lot of known-good ideas like residual connections. (ie, not so vanilla after all.) They end up needing a lot more parameters and flops to match resnet quality...

(Edit to add) Like why not use some clever distillation to get things to work well? Take the large well performing model and try to match intermediate activations to avoid gradient collapse. Training should then be very straightforward.

Really any time I see multi stage training in a paper I'm afraid... Tends to indicate that the whole system is very fragile, and that we'll be searching for just the right hyperparameters for months...

I would have liked to see it compared to MobileOne https://github.com/apple/ml-mobileone, which appears to run faster on an IPhone 12 Pro for comparable accuracy than VanillaNet does on an A100 server-class GPU.

It would also be great if there were pretrained mobiles available that better match typical mobile scenarios, e.g, comparable in complexity to Mobilenet2 0.5.

Transformers don't seem to optimal as an architecture for vision tasks, but I don't think pure convnets are either. Either way, having the right data is 10x as important.