about
Any Deep ReLU Network Is Shallow (arxiv.org)
18 points by fofoz on Jun 22, 2023 | hide | past | pdf | 5 comments on HN

In plain words: Any deep network built from ReLU units can be rewritten exactly as a three-layer one, with some weights allowed to be infinite; an algorithm finds those weights. The shallow version is transparent, so it can explain the original model's decisions.

Abstract · Any Deep ReLU Network is Shallow

We constructively prove that every deep ReLU network can be rewritten as a functionally identical three-layer network with weights valued in the extended reals. Based on this proof, we provide an algorithm that, given a deep ReLU network, finds the explicit weights of the corresponding shallow network. The resulting shallow network is transparent and used to generate explanations of the model s behaviour.

Mattia Jacopo Villani, Nandi Schoots
arXiv:2306.11827 · cs.LG, cs.AI, stat.ML · submitted Jun 20, 2023
abstract · pdf · html · 12 pages including bibliography and appendix

add comment on HN
Also discussed: Jun 2023 (134 points, 58 comments)

It doesn’t surprise me. It’s been known a long time that you can model arbitrary functions with a 3-layer network with logistic activation.
Do you have any further information or a source for this? As someone unfamiliar with ML, this sounds crazy to me.
It was first proposed here[1]: "Approximation by superpositions of a sigmoidal function"

[1]: https://link.springer.com/content/pdf/10.1007/BF02551274.pdf

It seems intuitive since ReLU is just a type of implicit regularization. Why would subsequent gradient descent help once you've achieved the benefit of throwing away the "outliers" or data beyond the threshold you want?
don't confuse this with universal approximation - yes shallow ReLU networks are dense in functional space, so at the limit you should be able to get any function you want - but they are talking about exact representation with finitely many neurons here.