about
The Cost of Training NLP Models: A Concise Overview (arxiv.org)
41 points by jonbaer on Apr 24, 2020 | hide | past | pdf | 5 comments on HN

In plain words: A review of what it costs to train large language models and which choices drive the bill, aimed at people planning experiments or trying to understand the economics. It breaks one intimidating price tag into the parts that matter for budgeting.

Abstract

We review the cost of training large-scale language models, and the drivers of these costs. The intended audience includes engineers and scientists budgeting their model-training experiments, as well as non-practitioners trying to make sense of the economics of modern-day Natural Language Processing (NLP).

Or Sharir, Barak Peleg, Yoav Shoham
arXiv:2004.08900 · cs.CL, cs.LG, cs.NE · submitted Apr 19, 2020
abstract · pdf · html

add comment on HN
Also discussed: Apr 2020 (2 points, 0 comments) · Apr 2020 (3 points, 0 comments) · Apr 2020 (5 points, 0 comments)

I don't understand how the conversion from FLOPs to dollars is done. A footnote says that "These $ figures come with substantial error bars, but we believe they are in the right ballpark." Where are the estimates themselves coming from? Cost of cloud computing? Cooling and power consumption estimates? I feel the article would benefit from adding a sentence fragment in the abstract or introduction that specifies this as I have no idea what I'm looking at right now.
I think the biggest advancement that could come out of the ML space right now is the equivalent of Big O notation for ML. Our current algorithms suck, and the first step to making them suck less is to precisely measure just how much they suck.
It’s similar to our viewpoint as well. Currently, it’s easy to see gains because we’re mainly going from highly manual processes to 1st generation AI applications. Future generation AI applications will mostly be optimization and squeezing out very tiny performance gains. We want to be able to highlight eventually how much things will cost and whether it’s even worth spending so much for a tiny amount of gain.

There’s so much hype right now that silly money is just being thrown at new tech without fully understanding the value you’re getting.

i mean for many classic ML models you have this — literally big O notation for the number of samples required to learn, number of iterations to convergence, cost per iteration, etc. You even get the cases where the algorithm with the best big-O bound is totally impractical due to constant factors.

The problem is just that all this goes out the window for deep learning, because nobody really knows how to prove much of anything about neural networks.

We have these for many models and parts of models. On one side, you have sample complexity [0] (how much data do you need?). On the other side, in deep learning you also have optimization convergence rates (how many steps do you need find the right parameters?). For example, newton's method converges at a quadratic rate for many problems. I know people work on similar theories for non-convex optimization, but the actual work there is outside what I know.

It would be great to have a breakthrough there.

[0] https://en.wikipedia.org/wiki/Sample_complexity