about
CCPS: Calibrating LLM Confidence via Perturbation Stability – EMNLP 2025 (arxiv.org)
3 points by erfan_mhi on Aug 28, 2025 | hide | past | pdf | 1 comment on HN

In plain words: It nudges a model's internal answer states, watches how much the output shifts, and uses a small classifier to judge if the answer is right. Across four models and three test variants, it cut calibration error by about 55% versus the best earlier method.

Abstract · Calibrating LLM Confidence by Probing Perturbed Representation Stability

Miscalibration in Large Language Models (LLMs) undermines their reliability, highlighting the need for accurate confidence estimation. We introduce CCPS (Calibrating LLM Confidence by Probing Perturbed Representation Stability), a novel method analyzing internal representational stability in LLMs. CCPS applies targeted adversarial perturbations to final hidden states, extracts features reflecting the model's response to these perturbations, and uses a lightweight classifier to predict answer correctness. CCPS was evaluated on LLMs from 8B to 32B parameters (covering Llama, Qwen, and Mistral architectures) using MMLU and MMLU-Pro benchmarks in both multiple-choice and open-ended formats. Our results show that CCPS significantly outperforms current approaches. Across four LLMs and three MMLU variants, CCPS reduces Expected Calibration Error by approximately 55% and Brier score by 21%, while increasing accuracy by 5 percentage points, Area Under the Precision-Recall Curve by 4 percentage points, and Area Under the Receiver Operating Characteristic Curve by 6 percentage points, all relative to the strongest prior method. CCPS delivers an efficient, broadly applicable, and more accurate solution for estimating LLM confidence, thereby improving their trustworthiness.

Reza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur, Ivan Brugere, Charese H. Smiley, Kundan Thind, Mohammad M. Ghassemi
arXiv:2505.21772 · cs.CL · submitted May 27, 2025 · updated Sep 18, 2025
abstract · pdf · html

add comment on HN

Author here. Our paper “Calibrating LLM Confidence by Probing Perturbed Representation Stability” was accepted to EMNLP 2025 Main Conference (top 15%) with a final rating of 9 (strong accept).

High-level summary: We probe LLM hidden states with slight perturbations to check answer stability—stable implies confidence; unstable implies uncertainty. This lightweight method delivers >50% reductions in calibration error (down to ~4.5%) across LLaMA, Mistral, Qwen on MMLU & MMLU-Pro, with no LLM fine-tuning.

Results, code, and dataset are available at: - Code: https://github.com/ledengary/CCPS - Data: https://huggingface.co/datasets/ledengary/CCPS

Happy to discuss technical details or calibration deployment strategies.