about
Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas (arxiv.org)
1 point by Brajeshwar 145 days ago | hide | past | pdf | discuss on HN

In plain words: Tested 33 AI models on 1,500 exam questions across six subject areas, checking how well their stated confidence matched which answers they got right. Scores varied by subject — applied knowledge was easiest to judge, formal reasoning and natural science hardest — hiding behind one overall number.

Abstract

Aggregate metacognitive quality scores mask within-model variation across MMLU benchmark domains. We administered 1,500 MMLU items (250 per domain, under an a priori six-domain grouping) to 33 frontier LLMs from eight model families and computed Type-2 AUROC per model-domain cell using verbalized confidence (0-100). Total observations: 47,151. Every model with above-chance aggregate monitoring showed non-trivial domain-level variation. Applied/Professional knowledge was reliably the easiest benchmark domain to monitor (mean AUROC = .742, ranked top-2 in 21 of 33 models); Formal Reasoning and Natural Science were reliably the hardest (one of the two ranked bottom-2 in 27 of 33 models). The three middle domains were statistically indistinguishable (Kendall's W = .164). A subject-level coherence analysis (within-domain similarity ratio = 0.95) confirms the six-domain grouping is a pragmatic benchmark taxonomy, not a validated latent construct. Within-family profile-shape clustering is significant for Anthropic, Google-Gemini, and Qwen (permutation p < .0001) but not DeepSeek, Google-Gemma, or OpenAI. Gemma 4 31B showed a +.202 AUROC improvement over Gemma 3 27B. Three models classified Invalid on binary KEEP/WITHDRAW probes produced normal profiles under verbalized confidence, confirming probe-format specificity. Bootstrap 95% CIs on 198 cells have median width .199. Split-half aggregate stability r = .893; profile-level split-half is weaker (grand median r = .184). These results show stable benchmark-domain variation obscured by aggregate metrics, and support benchmark-stage domain screening as a step before deployment in specific application areas.

Jon-Paul Cacioli
arXiv:2605.06673 · cs.CL, cs.AI, cs.LG · submitted Apr 21, 2026
abstract · pdf · html · 25 pages, 7 figures, 1 supplementary table. Code and data: https://github.com/synthiumjp/metacognitive-profile-atlas

add comment on HN