about
V1: Unifying Generation and Self-Verification for Parallel Reasoners (ArXiv) (arxiv.org)
2 points by harman2607 212 days ago | hide | past | pdf | 1 comment on HN

In plain words: Instead of scoring each candidate answer on its own, this system compares answers in pairs and spends its checking budget on the closest calls, with one model doing both solving and judging. It raised top-choice accuracy up to 10% over scoring answers separately.

Abstract · $V_1$: Unifying Generation and Self-Verification for Parallel Reasoners

Test-time scaling for complex reasoning tasks shows that leveraging inference-time compute, by methods such as independently sampling and aggregating multiple solutions, results in significantly better task outcomes. However, a critical bottleneck is verification: sampling is only effective if correct solutions can be reliably identified among candidates. While existing approaches typically evaluate candidates independently via scalar scoring, we demonstrate that models are substantially stronger at pairwise self-verification. Leveraging this insight, we introduce $V_1$, a framework that unifies generation and verification through efficient pairwise ranking. $V_1$ comprises two components: $V_1$-Infer, an uncertainty-guided algorithm using a tournament-based ranking that dynamically allocates self-verification compute to candidate pairs whose relative correctness is most uncertain; and $V_1$-PairRL, an RL framework that jointly trains a single model as both generator and pairwise self-verifier, ensuring the verifier adapts to the generator's evolving distribution. On code generation (LiveCodeBench, CodeContests, SWE-Bench) and math reasoning (AIME, HMMT) benchmarks, $V_1$-Infer improves Pass@1 by up to $10%$ over pointwise verification and outperforms recent test-time scaling methods while being significantly more efficient. Furthermore, $V_1$-PairRL achieves $7$--$9%$ test-time scaling gains over standard RL and pointwise joint training, and improves base Pass@1 by up to 8.7% over standard RL in a code-generation setting.

Harman Singh, Xiuyu Li, Kusha Sareen, Monishwaran Maheswaran, Sijun Tan, Xiaoxia Wu, Junxiong Wang, Alpay Ariyak, Qingyang Wu, Samir Khaki, Rishabh Tiwari, Long Lian, et al.
arXiv:2603.04304 · cs.CL · submitted Mar 4, 2026
abstract · pdf · html

add comment on HN

Hi HN, I’m one of the authors.

This paper studies how LLMs self-verify candidate solutions when doing test-time scaling (parallel reasoning / Best-of-N style generation).

We found that models are often much better at pairwise comparisons, where the model scores two solutions (A and B) jointly, than at assigning absolute scores to their own solutions independently.

The paper introduces:

• Pairwise self-verification instead of pointwise scoring

• V1-Infer, a ranking algorithm that selects good candidates efficiently

• V1-PairRL, RL training where generation and verification co-evolve to produce stronger self-verifiers

Across coding and reasoning benchmarks, we observe improved verification accuracy and good scaling when increasing verification compute budget.

One motivation is that many recent test-time scaling approaches (for example RSA: https://arxiv.org/abs/2509.26626 ) rely on sequential aggregation loops. Pairwise verification enables a more parallel form of selection, which may reduce latency in deep thinking pipelines and scaffolds.

Happy to answer questions.