about
Phi-4 Technical Report (arxiv.org)
2 points by veryluckyxyz on Dec 25, 2024 | hide | past | pdf | discuss on HN

In plain words: A language model with 14 billion adjustable values was trained mostly on carefully made practice data instead of scraped web text, keeping the same basic design as its predecessor. It beat the much larger teacher model it learned from on science and math questions.

Abstract

We present phi-4, a 14-billion parameter language model developed with a training recipe that is centrally focused on data quality. Unlike most language models, where pre-training is based primarily on organic data sources such as web content or code, phi-4 strategically incorporates synthetic data throughout the training process. While previous models in the Phi family largely distill the capabilities of a teacher model (specifically GPT-4), phi-4 substantially surpasses its teacher model on STEM-focused QA capabilities, giving evidence that our data-generation and post-training techniques go beyond distillation. Despite minimal changes to the phi-3 architecture, phi-4 achieves strong performance relative to its size -- especially on reasoning-focused benchmarks -- due to improved data, training curriculum, and innovations in the post-training scheme.

Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauffmann, James R. Lee, Yin Tat Lee, et al.
arXiv:2412.08905 · cs.CL, cs.AI · submitted Dec 12, 2024
abstract · pdf · html

add comment on HN
Also discussed: Dec 2024 (1 point, 0 comments) · Dec 2024 (3 points, 0 comments)