about
Pretraining on test set is all you need (arxiv.org)
3 points by nothrowaways on Oct 27, 2023 | hide | past | pdf | 1 comment on HN

In plain words: A tiny language model was trained on a dataset built only from the answers to academic benchmarks, essentially studying the test before taking it. It scored perfectly on those benchmarks, beating far larger models — but only because the test itself was the training data.

Abstract · Pretraining on the Test Set Is All You Need

Inspired by recent work demonstrating the promise of smaller Transformer-based language models pretrained on carefully curated data, we supercharge such approaches by investing heavily in curating a novel, high quality, non-synthetic data mixture based solely on evaluation benchmarks. Using our novel dataset mixture consisting of less than 100 thousand tokens, we pretrain a 1 million parameter transformer-based LLM \textbf{phi-CTNL} (pronounced ``fictional") that achieves perfect results across diverse academic benchmarks, strictly outperforming all known foundation models. \textbf{phi-CTNL} also beats power-law scaling and exhibits a never-before-seen grokking-like ability to accurately predict downstream evaluation benchmarks' canaries.

Rylan Schaeffer
arXiv:2309.08632 · cs.CL, cs.AI · submitted Sep 13, 2023
abstract · pdf · html · 3 pages, satire

add comment on HN
Also discussed: Aug 2025 (1 point, 0 comments) · Feb 2025 (2 points, 0 comments) · Oct 2023 (68 points, 26 comments) · Oct 2023 (4 points, 0 comments)

Heh, its satire... but they could probably get a little funding with some wording changes.

See also Brainderp, which actually does decently well for a 13B in the Open LLM leaderboard: https://huggingface.co/Sao10K/BrainDerp3