about
Fast Feedforward Networks (arxiv.org)
2 points by jasondavies on Sep 20, 2023 | hide | past | pdf | 1 comment on HN

In plain words: Instead of running every neuron in a layer, it sends each input down a decision tree to a few, so bigger layers cost barely more. It ran up to 220x faster than normal layers, keeping 94.2% of image accuracy with 1% of neurons.

Abstract

We break the linear link between the layer size and its inference cost by introducing the fast feedforward (FFF) architecture, a log-time alternative to feedforward networks. We demonstrate that FFFs are up to 220x faster than feedforward networks, up to 6x faster than mixture-of-experts networks, and exhibit better training properties than mixtures of experts thanks to noiseless conditional execution. Pushing FFFs to the limit, we show that they can use as little as 1% of layer neurons for inference in vision transformers while preserving 94.2% of predictive performance.

Peter Belcak, Roger Wattenhofer
arXiv:2308.14711 · cs.LG, cs.AI, cs.PF · submitted Aug 28, 2023 · updated Sep 18, 2023
abstract · pdf · html · 12 pages, 6 figures, 4 tables

add comment on HN
Also discussed: Sep 2023 (2 points, 0 comments)

This is a really cool idea. By dividing the space into regions, one avoids computing with neurons that do not play any role beside the given region.

This could be a game-changer in architectures like MLP-mixer and perhaps even transformers.