about
Two sparsities are better than one: Performance of sparse-sparse networks (arxiv.org)
51 points by Anon84 on Feb 11, 2022 | hide | past | pdf | 7 comments on HN

In plain words: A technique called Complementary Sparsity lines up a network's sparse connections and sparse active neurons so ordinary chips can process them efficiently. On FPGAs it delivered up to 100 times better throughput and energy efficiency for vision networks.

Abstract · Two Sparsities Are Better Than One: Unlocking the Performance Benefits of Sparse-Sparse Networks

In principle, sparse neural networks should be significantly more efficient than traditional dense networks. Neurons in the brain exhibit two types of sparsity; they are sparsely interconnected and sparsely active. These two types of sparsity, called weight sparsity and activation sparsity, when combined, offer the potential to reduce the computational cost of neural networks by two orders of magnitude. Despite this potential, today's neural networks deliver only modest performance benefits using just weight sparsity, because traditional computing hardware cannot efficiently process sparse networks. In this article we introduce Complementary Sparsity, a novel technique that significantly improves the performance of dual sparse networks on existing hardware. We demonstrate that we can achieve high performance running weight-sparse networks, and we can multiply those speedups by incorporating activation sparsity. Using Complementary Sparsity, we show up to 100X improvement in throughput and energy efficiency performing inference on FPGAs. We analyze scalability and resource tradeoffs for a variety of kernels typical of commercial convolutional networks such as ResNet-50 and MobileNetV2. Our results with Complementary Sparsity suggest that weight plus activation sparsity can be a potent combination for efficiently scaling future AI models.

Kevin Lee Hunter, Lawrence Spracklen, Subutai Ahmad
arXiv:2112.13896 · cs.LG, cs.AI, cs.AR, cs.NE · submitted Dec 27, 2021
abstract · pdf · html · 32 pages and 20 figures

add comment on HN

Always cool to see research from Numenta. They don’t get as much love as DeepMind, Google Brain, and OpenAI because their results aren’t as flashy, but I do feel like they’ve got a principled approach to engineering intelligent systems distinct from that of the big players.
Sparse operations still don’t have good support in any of the generic SIMD instructions (avx2/512, Neon), or in GPU…

And memory access/cache should also be optimized for them to have even more power saving and better speed.

For ASICs/FPGAs implementations, however, there’s no such limit and perf gains are higher.

> Sparse operations still don’t have good support in any of the generic SIMD instructions (avx2/512, Neon), or in GPU…

This is incorrect.

For GPUs: Nvidia A100 (launched in 2020) has sparse matrix multiplication support for inference.

For CPUs: Check out SLIDE from 2019: https://arxiv.org/pdf/1903.03129.pdf

Sparsity could really be a good thing but so many times I've tried to use it and walked away disappointed in terms of accuracy.
It is already a good thing, but it currently requires a lot of engineering effort to actually get it to work with acceptable quality. It's not something that works out of the box like half precision, or for some models, int8. And to your point, for many production scenarios the ratio of engineering work vs performance gains is maybe not worth it. But for models that are going to handle massive load in inference it is worth it in my experience.

I expect that this will be made much easier in future with better hardware support and smarter sparsification libraries.

It's the power savings that are the real goal. If you can create and (re)train a network for incredible sparsity you can fit inference in some pretty low power envelopes. I think to that end the work required to get to adequate perform is justified
An abundance of sparsities.