In plain words: Simply writing the same prompt more than once before asking for an answer makes popular chat models answer better when they are not thinking step by step. It beats the usual single prompt without generating extra words or taking longer.
Abstract
When not using reasoning, repeating the input prompt improves performance for popular models (Gemini, GPT, Claude, and Deepseek) without increasing the number of generated tokens or latency.
Yaniv Leviathan, Matan Kalman, Yossi Matias
arXiv:2512.14982 · cs.LG, cs.AI, cs.CL · submitted Dec 17, 2025
abstract · pdf · html
The question I can't currently answer: how much of the benefit comes from the semantic content of the axioms versus the repetition/emphasis effect this paper identifies?
I'm running an ablation study with a critical condition: shuffled axioms (same tokens, randomized order). If shuffled matches structured axioms, the content doesn't matter. If structured axioms win, semantic structure genuinely helps beyond repetition.
I'll add this experiment to the parent project on OSF. Curious whether others working on knowledge injection techniques have similar confounds to untangle.