about
Biases in the Blind Spot: Detecting What LLMs Fail to Mention (arxiv.org)
3 points by mpweiher 225 days ago | hide | past | pdf | discuss on HN

In plain words: A tool guesses hidden biases, then tests each by changing one trait in a task's inputs and seeing if answers shift even though the model's reasoning never mentions it. On seven models judging hiring, loans, and admissions, it found unknown biases like Spanish fluency.

Abstract

Large Language Models (LLMs) often provide chain-of-thought (CoT) reasoning traces that appear plausible, but may hide internal biases. We call these unverbalized biases. Monitoring models via their stated reasoning is therefore unreliable, and existing bias evaluations typically require predefined categories and hand-crafted datasets. In this work, we introduce a fully automated, black-box pipeline for detecting task-specific unverbalized biases. Given a task dataset, the pipeline uses LLM autoraters to generate candidate bias concepts. It then tests each concept on progressively larger input samples by generating positive and negative variations, and applies statistical techniques for multiple testing and early stopping. A concept is flagged as an unverbalized bias if it yields statistically significant performance differences while not being cited as justification in the model's CoTs. We evaluate our pipeline across seven LLMs on three decision tasks (hiring, loan approval, and university admissions). Our technique automatically discovers previously unknown biases in these models (e.g., Spanish fluency, English proficiency, writing formality). In the same run, the pipeline also validates biases that were manually identified by prior work (gender, race, religion, ethnicity). More broadly, our proposed approach provides a practical, scalable path to automatic, more efficient, and broader task-specific unverbalized bias discovery.

Iván Arcuschin, David Chanin, Adrià Garriga-Alonso, Oana-Maria Camburu
arXiv:2602.10117 · cs.LG, cs.AI · submitted Feb 10, 2026 · updated May 29, 2026
abstract · pdf · html · Published at the 43rd International Conference on Machine Learning (ICML 2026)

add comment on HN
Also discussed: Feb 2026 (2 points, 0 comments) · Feb 2026 (2 points, 0 comments) · Feb 2026 (1 point, 0 comments) · Feb 2026 (4 points, 1 comment)