In plain words: Four simple puzzle-like tasks were built to test how standard deep learning training, which nudges weights step by step to cut errors, handles them. On all four, that usual training either fails outright or struggles badly, and the analysis explains why.
Abstract · Failures of Gradient-Based Deep Learning
In recent years, Deep Learning has become the go-to solution for a broad range of applications, often outperforming state-of-the-art. However, it is important, for both theoreticians and practitioners, to gain a deeper understanding of the difficulties and limitations associated with common approaches and algorithms. We describe four types of simple problems, for which the gradient-based algorithms commonly used in deep learning either fail or suffer from significant difficulties. We illustrate the failures through practical experiments, and provide theoretical insights explaining their source, and how they might be remedied.
Shai Shalev-Shwartz, Ohad Shamir, Shaked Shammah
arXiv:1703.07950 · cs.LG, cs.NE, stat.ML · submitted Mar 23, 2017 · updated Apr 26, 2017
abstract · pdf
E.g. page 5. They attempt to explain a really simple idea, that they generated images of random lines at a random angle. Then labelled the lines positive or negative examples, based on whether the angle was greater than 90 degrees or not. Then they take sets of these examples. And label them based on whether they contain an even or odd number of positive examples.
They take several paragraphs over half a page to explain this. Filled with dense mathematical notation. If you don't know what symbols like U, :, ->, or ~ mean, you are screwed because that's not googleable. It takes way longer to parse than it should. Especially since I just wanted to quickly skim the ideas, not painfully reverse engineer them.
Hell, even the concept of even or odd, is pointlessly redefined and complicated as multiplying + or - 1's together. I was scratching my head for a few minutes just trying to figure out what the purpose of that was. It's like reading bad code without any comments. Even if you are very familiar with the language and know what the code does, it takes a lot of effort to figure out why it's that way. If it's not explained properly.
The worst part is, no one ever complains about this stuff because they are afraid of looking stupid. I sure fear that by posting this very comment. I actually am familiar with the notation used in this example. I still find it unnecessary and exhausting to decode.