about
The Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge (arxiv.org)
1 point by sbulaev 82 days ago | hide | past | pdf | discuss on HN

In plain words: Judge bias shows up inside the model's internal signals as a specific direction; nudging activity along it shifts scores, and reversing it cancels the bias. A simple readout of that direction predicted judge failures on three new benchmarks better than text-based checks.

Abstract · Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operationally useful in ways it does not afford. We report three findings, across seven judges, seven bias types, and nine benchmarks. Geometry: baseline judging inputs occupy a tight activation manifold while biased inputs are displaced along a low-dimensional, type-specific subspace that sharpens with depth and is recovered consistently by three families of estimators. Causal control: steering hidden states along this subspace drives scoring in both directions, forward shifts reproducing biased scoring on clean inputs and reverse shifts restoring baseline scoring on biased ones, while matched-norm random directions produce shifts an order of magnitude smaller. Operational: a simple linear projection onto the same bias-direction features anticipates judge failures on three entirely unseen benchmarks, substantially outperforming text-based alternatives. Reading bias as activation geometry, rather than as input-output noise, unifies geometric structure, causal control, and operational prediction within a single framework. The project page is available at https://xzx34.github.io/unfair-judge/

Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li, Zirui Song, Xiuying Chen
arXiv:2607.11871 · cs.LG, cs.AI, cs.CL · submitted Jul 13, 2026
abstract · pdf · html · 58 pages, 13 figures, 30 tables; project page: https://xzx34.github.io/unfair-judge/

add comment on HN