In plain words: A cheap judge gives only the odds for each verdict, and its confidence decides whether to accept it or hand the task to a slower reasoning judge. On 1,610 held-out pairs this routing beat GPT-6 by 0.9 points at 41% of its cost.
Abstract · JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.
Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman
arXiv:2609.26550 · cs.AI · submitted Sep 22, 2026 · updated Sep 29, 2026
abstract · pdf · html · Expanded the dataset, updated the results and figures, and added new analyses. The previous result reporting 99% of GPT performance at 57% of the cost is retained in the appendix