In plain words: Shrinking a reasoning model's numbers to save memory makes it ramble longer; in up to 52% of mistakes it had the right answer but never said it. Penalizing words such as wait cuts rambling 12-23% without losing accuracy and reduces these mistakes up to 58%.
Abstract
Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood. Across math, coding, and science QA, we find that aggressive PTQ reduces accuracy while increasing chain-of-thought (CoT) length. Surprisingly, we show that in up to 52% of the quantized models' failures, models reach the right answer in intermediate reasoning steps but do not output it as a final answer. To understand why quantization leads to this increase in overthinking errors, we measure the token-level KL divergence between quantized and full-precision output distributions. Positions with high KL divergence correlate strongly with high next-token entropy, and at these positions quantized models disproportionately sample overthinking markers such as "wait", "but", and "alternatively". We show that simply introducing a training-free logit penalty on a curated set of overthinking markers can reduce CoT length by 12--23% while preserving or improving accuracy across 5 models (1.5B-32B parameters), 3 quantization methods, and 5 benchmarks, yielding a favorable Pareto frontier of accuracy against reasoning cost compared to penalizing other token sets. Overthinking errors produced by quantized models are particularly reduced by up to 58%.
Sanae Lotfi, Polina Kirichenko, Steven Li, Zechun Liu
arXiv:2606.00206 · cs.LG · submitted May 29, 2026
abstract · pdf · html
Highly quantized models, especially with highly quantized KV caches, will, effectively, attend to the wrong tokens and be unable to easily discern highly similar tokens. The bastardized way of explaining this is gradient descent techniques get stuck in localized minimum and global maximums, so what happens when you turn the slopes into hard stair steps?
We need to move to smaller models and smaller caches and better samplers, not new quant methods (although I'm willing to also take those too).