In plain words: A code-writing model was trained on web text filtered for textbook quality plus AI-made textbooks and exercises, instead of huge piles of raw code. Though far smaller than rivals, it solved 50.6% of test coding problems on the first try.
Abstract · Textbooks Are All You Need
We introduce phi-1, a new large language model for code, with significantly smaller size than competing models: phi-1 is a Transformer-based model with 1.3B parameters, trained for 4 days on 8 A100s, using a selection of ``textbook quality" data from the web (6B tokens) and synthetically generated textbooks and exercises with GPT-3.5 (1B tokens). Despite this small scale, phi-1 attains pass@1 accuracy 50.6% on HumanEval and 55.5% on MBPP. It also displays surprising emergent properties compared to phi-1-base, our model before our finetuning stage on a dataset of coding exercises, and phi-1-small, a smaller model with 350M parameters trained with the same pipeline as phi-1 that still achieves 45% on HumanEval.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, et al.
arXiv:2306.11644 · cs.CL, cs.AI, cs.LG · submitted Jun 20, 2023 · updated Oct 2, 2023
abstract · pdf · html · 26 pages; changed color scheme of plot. fixed minor typos and added couple clarifications
Now this wouldn't be possible without the high quality synthetic dataset produced by GPT(1B tokens) but this is more evidence in line with Tiny Stories (https://arxiv.org/abs/2305.07759). That is, LLMs only need to be so big (both data and parameters) to learn the total sum of human knowledge (and deal with trash data).