In plain words: A tiny language model was trained on a dataset built only from the answers to academic benchmarks, essentially studying the test before taking it. It scored perfectly on those benchmarks, beating far larger models — but only because the test itself was the training data.
Abstract
Inspired by recent work demonstrating the promise of smaller Transformer-based language models pretrained on carefully curated data, we supercharge such approaches by investing heavily in curating a novel, high quality, non-synthetic data mixture based solely on evaluation benchmarks. Using our novel dataset mixture consisting of less than 100 thousand tokens, we pretrain a 1 million parameter transformer-based LLM \textbf{phi-CTNL} (pronounced ``fictional") that achieves perfect results across diverse academic benchmarks, strictly outperforming all known foundation models. \textbf{phi-CTNL} also beats power-law scaling and exhibits a never-before-seen grokking-like ability to accurately predict downstream evaluation benchmarks' canaries.
Rylan Schaeffer
arXiv:2309.08632 · cs.CL, cs.AI · submitted Sep 13, 2023
abstract · pdf · html · 3 pages, satire
What's left, after you use, say, "half" of The Entire Internet for training? What's left that is definitely not the same as the half you used? How can you be sure, when you're working with terabytes or petabytes?
Credible benchmarking is going to have to become a field in itself.