about
On Efficient Training of Large-Scale Deep Learning Models: A Literature Review (arxiv.org)
38 points by PaulHoule on Apr 10, 2023 | hide | past | pdf | 3 comments on HN

In plain words: This review sorts ways to train deep-learning models faster by which part of the training step they change: the data, the model, the update rules, tight budgets, or the hardware. It explains how each cuts work and how they combine, where no guide existed.

Abstract

The field of deep learning has witnessed significant progress, particularly in computer vision (CV), natural language processing (NLP), and speech. The use of large-scale models trained on vast amounts of data holds immense promise for practical applications, enhancing industrial productivity and facilitating social development. With the increasing demands on computational capacity, though numerous studies have explored the efficient training, a comprehensive summarization on acceleration techniques of training deep learning models is still much anticipated. In this survey, we present a detailed review for training acceleration. We consider the fundamental update formulation and split its basic components into five main perspectives: (1) data-centric: including dataset regularization, data sampling, and data-centric curriculum learning techniques, which can significantly reduce the computational complexity of the data samples; (2) model-centric, including acceleration of basic modules, compression training, model initialization and model-centric curriculum learning techniques, which focus on accelerating the training via reducing the calculations on parameters; (3) optimization-centric, including the selection of learning rate, the employment of large batchsize, the designs of efficient objectives, and model average techniques, which pay attention to the training policy and improving the generality for the large-scale models; (4) budgeted training, including some distinctive acceleration methods on source-constrained situations; (5) system-centric, including some efficient open-source distributed libraries/systems which provide adequate hardware support for the implementation of acceleration algorithms. By presenting this comprehensive taxonomy, our survey presents a comprehensive review to understand the general mechanisms within each component and their joint interaction.

Li Shen, Yan Sun, Zhiyuan Yu, Liang Ding, Xinmei Tian, Dacheng Tao
arXiv:2304.03589 · cs.LG, cs.AI, cs.DC · submitted Apr 7, 2023
abstract · pdf · html · 60 pages

add comment on HN

See also "A Bibliometric Review of Large Language Models Research from 2017 to 2023"

https://arxiv.org/abs/2304.02020

The tooling on arxiv is great, but semanticscholar has some nice paper navigation features as well.

https://www.semanticscholar.org/paper/A-Bibliometric-Review-...

I hate the 2-d UMAP maps that always show those hyper-dimensional cusps though and think well-tuned tSNE is much more civilized. I was talking to the people at arXiv the other day and found out they are working on an HTML viewer for most papers that will make arXiv a much better target for linking.
I am not smart enough to get upset over UMAP vs tSNE. 3d bar charts, no error bars and unlabeled axis are still my windmills.

If you are talking to arXiv folks, please get them to fix author search, which is horribly broken for Asian names.

Take this paper for example, https://arxiv.org/abs/2304.03717 if I then click on author, "Yuanzhi Li" it does a search using "Li, Y" which reports 4768 results.

The search for the full name gives a much better 86 results.

https://arxiv.org/search/cs?query=Yuanzhi+Li&searchtype=auth...

Ideally, they would use the author's full quoted name in the search.

https://arxiv.org/search/cs?query=%22Yuanzhi+Li%22&searchtyp...

I have reported this multiple times, nothing.