In plain words: Models can notice a signal pushed into their inner workings and name the concept; an early-layer circuit spots it and turns off the default "no". Only preference training, not ordinary fine-tuning, creates it; a trained bias nudge raised detection 75% without more false alarms.
Abstract · Mechanisms of Introspective Awareness
Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introspective awareness." We investigate the mechanisms underlying this capability in open-weights models. First, we find that it is behaviorally robust: models detect injected steering vectors at moderate rates with 0% false positives across diverse prompts and dialogue formats. Notably, this capability emerges specifically from post-training; we show that preference optimization algorithms like DPO can elicit it, but standard supervised finetuning does not. We provide evidence that detection cannot be explained by simple linear association between certain steering vectors and directions promoting affirmative responses. We trace the detection mechanism to a two-stage circuit in which "evidence carrier" features in early post-injection layers detect perturbations monotonically along diverse directions, suppressing downstream "gate" features that implement a default negative response. This circuit is absent in base models and robust to refusal ablation. Identification of injected concepts relies on largely distinct later-layer mechanisms that only weakly overlap with those involved in detection. Finally, we show that introspective capability is substantially underelicited: ablating refusal directions improves detection by +53%, and a trained bias vector improves it by +75% on held-out concepts, both without meaningfully increasing false positives. Our results suggest that this introspective awareness of injected concepts is robust and mechanistically nontrivial, and could be substantially amplified in future models. Code: https://github.com/safety-research/introspection-mechanisms.
Uzay Macar, Li Yang, Atticus Wang, Peter Wallich, Emmanuel Ameisen, Jack Lindsey
arXiv:2603.21396 · cs.LG · submitted Mar 22, 2026 · updated Jun 10, 2026
abstract · pdf · html
If I understood the paper right, during post-training (e.g. DPO) models learn the correct "shape" of responses. But unlike SFT they're also penalized for going off-manifold so this incentivizes development of circuits that can detect off-manifold responses (you can see this clearly with RLVR perhaps, where models have a "but wait" reflex to steer themselves back in the correct direction) [^1]. Since part of the training is to be the archetypical chatbot assistant though, when combining with anti-jailbreak training this usually gets linked into "refusal" circuits.
One hypothesis might be that the question itself is leading. I.e. models will by default respond "no" to "are there any injected thoughts", just as they would to "are you conscious" or "do you have feeligns", because of RLHF that triggers refusal behavior. Then injection provides a strong enough signal that ends up "scrambling" this pathway, _suppressing_ the normal refusal behavior and allowing them to report the injection. (Describing the contents of the injected vector is trivial either way, as the paper notes the detection is the important part).
The interesting thing is that ablating away refusals doesn't actually change the false positive rate though, so instead the above hypothesis of injections overriding a default refusal doesn't fit. Instead there really does seem to be a separate "evidence carrier" detector sensitive to off-manifold responses, which just so happens to get wired into the "refusal circuits" but when "unwired" via ablation allows the model to report injections.
I guess what's not clear to me though is whether this is really detecting _injection_ itself. Wouldn't the same circuits be triggered by any anomalous context? It shouldn't be any surprise that models can detect models anomalies in input tokens (after all LLMs were designed to model text), so I don't see why anomalies in the residual stream would be any different (it's not like a layer cares whether the "bread" embedding was injected externally or came through from the input token).
In theory the case of "anomalous input context" versus "anomalous residual via external injection" _can_ be distinguished though, because there would be a sort of "discontinuity" in the residual stream as you pass through layers, and the hidden state at token i depth n feeds into that of token i+1 depth n+1, you could in theory create a computational graph that could detect such tampering.
I think the paper sort of indirectly tested this in section 3.2 "SPECIFICITY TO THE ASSISTANT PERSONA"
>In contrast, the two nonstandard roles (Alice-Bob, story framing) induce confabulation. Thus, introspection is not exclusive to responding as the assistant character, although reliability decreases outside standard roles.
Which does seem to imply that as soon as you step out of distribution to things like roleplay that RLHF specifically penalized, the anomaly detectors start firing as well.
[^1] I think this is also related to how RLHF/DPO are sequence-level optimizations, with a notion of credit assignment. And optimizing in this way results in the model having a notion of whether the current position in the rollout is "good" or not.