about
Distillation Scaling Laws (arxiv.org)
5 points by brandonb on Aug 15, 2025 | hide | past | pdf | discuss on HN

In plain words: A formula predicts how well a student learns from a teacher's answers, based on how the computing budget is split. Copying beats standard training when a teacher already exists or many students need it; standard training wins if you train a teacher for one student.

Abstract

We propose a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings mitigate the risks associated with large-scale distillation by enabling compute-optimal allocation for both the teacher and student to maximize student performance. We provide compute-optimal distillation recipes for two key scenarios: when a teacher already exists, and when a teacher needs training. In settings involving many students or an existing teacher, distillation outperforms supervised learning up to a compute level that scales predictably with student size. Conversely, if only one student is to be distilled and a teacher also requires training, supervised learning is generally preferable. Additionally, our large-scale study of distillation increases our understanding of the process and helps inform experimental design.

Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russ Webb
arXiv:2502.08606 · cs.LG, cs.AI, cs.CL, stat.ML · submitted Feb 12, 2025 · updated Jul 25, 2025
abstract · pdf · html · Version accepted to ICML 2025. 69 pages, 54 figures, 13 tables

add comment on HN
Also discussed: Feb 2025 (3 points, 0 comments)