about
Energy and Carbon Considerations of Fine-Tuning Bert (arxiv.org)
1 point by PaulHoule on Nov 25, 2023 | hide | past | pdf | discuss on HN

In plain words: They measured the energy and carbon used to fine-tune a language model across tasks, data, and computers, since past estimates counted only the initial training step. Fine-tuning is far cheaper per run, but done so often by many people that its total footprint matters.

Abstract · Energy and Carbon Considerations of Fine-Tuning BERT

Despite the popularity of the `pre-train then fine-tune' paradigm in the NLP community, existing work quantifying energy costs and associated carbon emissions has largely focused on language model pre-training. Although a single pre-training run draws substantially more energy than fine-tuning, fine-tuning is performed more frequently by many more individual actors, and thus must be accounted for when considering the energy and carbon footprint of NLP. In order to better characterize the role of fine-tuning in the landscape of energy and carbon emissions in NLP, we perform a careful empirical study of the computational costs of fine-tuning across tasks, datasets, hardware infrastructure and measurement modalities. Our experimental results allow us to place fine-tuning energy and carbon costs into perspective with respect to pre-training and inference, and outline recommendations to NLP researchers and practitioners who wish to improve their fine-tuning energy efficiency.

Xiaorong Wang, Clara Na, Emma Strubell, Sorelle Friedler, Sasha Luccioni
arXiv:2311.10267 · cs.CL, cs.LG · submitted Nov 17, 2023 · updated Oct 16, 2024
abstract · pdf · html · EMNLP 2023 Findings; First two authors contributed equally; 12 pages

add comment on HN