In plain words: A small helper model drafts answers and a reward model scores them, so the big model steers toward high-reward replies without generating many full answers. It beat the usual trick of generating many answers and picking the best, cutting latency by up to 28%.
Abstract · Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of-$n$ test-time scaling with a reward model $r(x,y)$ and speculative samples from a small auxiliary model $π_S(y\mid x)$. We provably approximate both the optimal tilted policy $π_{β,B}(y\mid x) \propto π_B(y\mid x)\exp(β\,r(x,y))$ of soft best-of-$n$ under the base model $π_B$, as well as the expected reward under the optimal policy. In experiments on reasoning benchmarks (MATH500, OlympiadBench, Minerva Math, MMLU-STEM, GSM8K) and across different model families, our method achieves higher accuracy than standard soft best-of-$n$ with $π_S$ and reward-guided speculative decoding (Liao et al., 2025), and in certain settings even outperforms soft best-of-$n$ with $π_B$, while reducing end-to-end latency by up to $28\%$. The code is available at https://github.com/j-geuter/GSI .
Jonathan Geuter, Youssef Mroueh, David Alvarez-Melis
arXiv:2506.04118 · cs.LG, stat.ML · submitted Jun 4, 2025 · updated Apr 27, 2026
abstract · pdf · html · 41 pages, 11 figures. Published at ICLR 2026