In plain words: Neural networks are essentially polynomial regression in disguise, fitting curves by multiplying input features together, which explains their quirks. In tests, the plain polynomial version matched or beat neural nets while needing fewer tuning choices and avoiding training failures.
Abstract · Polynomial Regression As an Alternative to Neural Nets
Despite the success of neural networks (NNs), there is still a concern among many over their "black box" nature. Why do they work? Here we present a simple analytic argument that NNs are in fact essentially polynomial regression models. This view will have various implications for NNs, e.g. providing an explanation for why convergence problems arise in NNs, and it gives rough guidance on avoiding overfitting. In addition, we use this phenomenon to predict and confirm a multicollinearity property of NNs not previously reported in the literature. Most importantly, given this loose correspondence, one may choose to routinely use polynomial models instead of NNs, thus avoiding some major problems of the latter, such as having to set many tuning parameters and dealing with convergence issues. We present a number of empirical results; in each case, the accuracy of the polynomial approach matches or exceeds that of NN approaches. A many-featured, open-source software package, polyreg, is available.
Xi Cheng, Bohdan Khomtchouk, Norman Matloff, Pete Mohanty
arXiv:1806.06850 · cs.LG, stat.ML · submitted Jun 13, 2018 · updated Apr 10, 2019
abstract · pdf · html · 23 pages, 1 figure, 13 tables
No good ML practitioner believes that tiny, shallow, fully-connected neural networks are the best algorithm for every problem (just look at Kaggle results if you don't believe me). Especially for small, easy datasets like the ones analyzed in this paper, NNs are often not the best choice. However, for large scale image classification, density modeling, sequence modeling, etc. (none of which are tested in the paper), NNs are SOTA.
Hilariously, the paper only compares polynomial regression to tiny neural networks. I bet if they had thrown in results from XGBoost or other classical ML techniques, polynomial regression would be blown out of the water.
There's also some heuristics you can use to tell when a paper like this isn't necessarily representative of the field of ML:
- Missing citations (e.g. "It is well-known that NNs are prone to overfitting [Chollet and Allaire(2018)], which has been the subject of much study, e.g. [?].").
- Inconsistent formatting (e.g. the tables on page 7)
- Sentences like "Much more empirical work is needed to explore these issues." following something that sounds like an easy experiment to try.