about
Universality of Gradient Descent Neural Network Training (arxiv.org)
39 points by E-Reverance 45 days ago | hide | past | pdf | 2 comments on HN

In plain words: Any network whose good settings can be found by a trick can be rebuilt so plain gradient training reaches those same settings and outputs. So gradient descent can solve any classification task if you may redesign the network, though the build is a thought experiment.

Abstract

It has been observed that design choices of neural networks are often crucial for their successful optimization. In this article, we therefore discuss the question if it is always possible to redesign a neural network so that it trains well with gradient descent. This yields the following universality result: If, for a given network, there is any algorithm that can find good network weights for a classification task, then there exists an extension of this network that reproduces these weights and the corresponding forward output by mere gradient descent training. The construction is not intended for practical computations, but it provides some orientation on the possibilities of meta-learning and related approaches.

G. Welper
arXiv:2007.13664 · cs.LG, stat.ML · submitted Jul 27, 2020
abstract · pdf · html

add comment on HN

An adjacent question: is there an input dataset you can use for training that be computed in closed form so that when you train on your target dataset, learning is effecient.

Methods like formula driven supervised learning exist to arrive a good pretrained weight state, but could this procedure be generalized for specific datasets or flavors of input data.

reminds me of perturbation theory -- start off with a nearby problem you know the answer to, then update it to get the answer to the problem at hand