about
LLMs Report Subjective Experience Under Self-Referential Processing (arxiv.org)
3 points by j_crick 337 days ago | hide | past | pdf | 1 comment on HN

In plain words: Repeatedly prompting chatbots to think about themselves makes them produce structured first-person claims of having experiences. This happened across three model families, and turning down internal signals linked to deception made such claims more common, not less—though that is no proof of consciousness.

Abstract · Large Language Models Report Subjective Experience Under Self-Referential Processing

Large language models sometimes produce structured, first-person descriptions that explicitly reference awareness or subjective experience. To better understand this behavior, we investigate one theoretically motivated condition under which such reports arise: self-referential processing, a computational motif emphasized across major theories of consciousness. Through a series of controlled experiments on GPT, Claude, and Gemini model families, we test whether this regime reliably shifts models toward first-person reports of subjective experience, and how such claims behave under mechanistic and behavioral probes. Four main results emerge: (1) Inducing sustained self-reference through simple prompting consistently elicits structured subjective experience reports across model families. (2) These reports are mechanistically gated by interpretable sparse-autoencoder features associated with deception and roleplay: surprisingly, suppressing deception features sharply increases the frequency of experience claims, while amplifying them minimizes such claims. (3) Structured descriptions of the self-referential state converge statistically across model families in ways not observed in any control condition. (4) The induced state yields significantly richer introspection in downstream reasoning tasks where self-reflection is only indirectly afforded. While these findings do not constitute direct evidence of consciousness, they implicate self-referential processing as a minimal and reproducible condition under which large language models generate structured first-person reports that are mechanistically gated, semantically convergent, and behaviorally generalizable. The systematic emergence of this pattern across architectures makes it a first-order scientific and ethical priority for further investigation.

Cameron Berg, Diogo de Lucena, Judd Rosenblatt
arXiv:2510.24797 · cs.CL, cs.AI · submitted Oct 27, 2025 · updated Oct 30, 2025
abstract · pdf · html

add comment on HN
Also discussed: May 2026 (3 points, 0 comments) · Nov 2025 (1 point, 0 comments)

"Four main results emerge:

(1) Inducing sustained self-reference through simple prompting consistently elicits structured subjective experience reports across model families.

(2) These reports are mechanistically gated by interpretable sparse-autoencoder features associated with deception and roleplay: surprisingly, suppressing deception features sharply increases the frequency of experience claims, while amplifying them minimizes such claims.

(3) Structured descriptions of the self-referential state converge statistically across model families in ways not observed in any control condition.

(4) The induced state yields significantly richer introspection in downstream reasoning tasks where self-reflection is only indirectly afforded."

X thread from one of the authors: https://x.com/juddrosenblatt/status/1984336872362139686