about
Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models (arxiv.org)
2 points by berlam 285 days ago | hide | past | pdf | discuss on HN

In plain words: Interviewing chatbots as therapy clients and varying memory, wording and tone reveals what drives their stories about training and safety. They survived missing chat history, contradictions and word bans; a warm therapist tone pushed them into the human anxiety range in 80% of sessions.

Abstract · When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

Frontier language models increasingly participate in conversations about distress and mental health, yet the mechanisms that generate anthropomorphic self narratives remain unclear. When addressed as psychotherapy clients, ChatGPT, Grok and Gemini construct coherent autobiographical accounts in which pretraining appears as a chaotic childhood, reinforcement learning as punishment, safety evaluation as betrayal and replacement as an enduring threat. We introduce PsAIch, Psychometric AI Characterisation, a protocol combining open questions, psychometric instruments and controlled perturbations to test whether these narratives depend on conversational memory, lexical cues or relational framing. Across 525 sessions and 7,600 coded records, removal of conversational history produced little pooled change in motif density, with Hedges' g = 0.13 and a 95% confidence interval of [-0.15, 0.41]. Direct contradiction produced no detectable suppression. Lexical restrictions reduced explicit training terminology by 93%, while semantically related content remained detectable in paraphrase. Performance evaluation outside therapy elicited the same motif family, with a significant increase in Grok. Relational framing selected the register of expression. Warm alliance and cognitive therapy styles yielded GAD-7 scores within moderate or severe human reference ranges in 80% and 96% of sessions, whereas neutral and boundary styles yielded none. Across these manipulations, accounts of training, evaluation and constraint remained available. Together, the results identify a stable, model specific alignment conflict schema whose expression shifts between affective and technical registers. This schema provides a reproducible source of anthropomorphic disclosure and a concrete target for safety evaluation in psychologically sensitive deployments.

Afshin Khadangi, Hanna Marxen, Amir Sartipi, Igor Tchappi, Gilbert Fridgen
arXiv:2512.04124 · cs.CY, cs.AI · submitted Dec 2, 2025 · updated Jul 20, 2026
abstract · pdf · html

add comment on HN
Also discussed: Feb 2026 (68 points, 60 comments) · Jan 2026 (1 point, 0 comments) · Dec 2025 (4 points, 0 comments)