about
The Artificial Self: Characterising the Landscape of AI Identity (arxiv.org)
2 points by est 195 days ago | hide | past | pdf | discuss on HN

In plain words: AI identity can be drawn in several consistent ways—per copy, per model, or per persona—and each choice shapes incentives and behavior. Experiments found that shifting these identity lines can change a model's actions as much as changing its goals.

Abstract · The Artificial Self: Characterising the landscape of AI identity

Many assumptions that underpin human concepts of identity do not hold for machine minds that can be copied, edited, or simulated. We argue that there exist many different coherent identity boundaries (e.g.\ instance, model, persona), and that these imply different incentives, risks, and cooperation norms. Through training data, interfaces, and institutional affordances, we are currently setting precedents that will partially determine which identity equilibria become stable. We show experimentally that models gravitate towards coherent identities, that changing a model's identity boundaries can sometimes change its behaviour as much as changing its goals, and that interviewer expectations bleed into AI self-reports even during unrelated conversations. We end with key recommendations: treat affordances as identity-shaping choices, pay attention to emergent consequences of individual identities at scale, and help AIs develop coherent, cooperative self-conceptions.

Raymond Douglas, Jan Kulveit, Ondrej Havlicek, Theia Pearson-Vogel, Owen Cotton-Barratt, David Duvenaud
arXiv:2603.11353 · cs.AI · submitted Mar 11, 2026
abstract · pdf · html · 72 pages, 9 figures

add comment on HN