about
Understanding LLMs: A Comprehensive Overview from Training to Inference (arxiv.org)
3 points by Anon84 on Jan 10, 2024 | hide | past | pdf | discuss on HN

In plain words: It traces how large language models are built and run to answer questions, from cleaning training data and teaching the model to shrinking it for cheaper hardware. The survey maps cost-cutting tricks at each stage, showing where training and everyday use can get cheaper.

Abstract

The introduction of ChatGPT has led to a significant increase in the utilization of Large Language Models (LLMs) for addressing downstream tasks. There's an increasing focus on cost-efficient training and deployment within this context. Low-cost training and deployment of LLMs represent the future development trend. This paper reviews the evolution of large language model training techniques and inference deployment technologies aligned with this emerging trend. The discussion on training includes various aspects, including data preprocessing, training architecture, pre-training tasks, parallel training, and relevant content related to model fine-tuning. On the inference side, the paper covers topics such as model compression, parallel computation, memory scheduling, and structural optimization. It also explores LLMs' utilization and provides insights into their future development.

Yiheng Liu, Hao He, Tianle Han, Xu Zhang, Mengyuan Liu, Jiaming Tian, Yutong Zhang, Jiaqi Wang, Xiaohui Gao, Tianyang Zhong, Yi Pan, Shaochen Xu, et al.
arXiv:2401.02038 · cs.CL · submitted Jan 4, 2024 · updated Jan 6, 2024
abstract · pdf · html · 30 pages,6 figures

add comment on HN
Also discussed: Feb 2024 (2 points, 0 comments)