about
What makes your model a low-empathy or warmth person? (arxiv.org)
1 point by PaulHoule on Oct 23, 2024 | hide | past | pdf | discuss on HN

In plain words: They hunt for hidden signals inside a language model that stand for things like cultural norms and stress, then nudge those signals to shift its personality without any retraining. This shows how such factors shape the model's personality and, through it, its safety.

Abstract · Exploring the Personality Traits of LLMs through Latent Features Steering

Large language models (LLMs) have significantly advanced dialogue systems and role-playing agents through their ability to generate human-like text. While prior studies have shown that LLMs can exhibit distinct and consistent personalities, the mechanisms through which these models encode and express specific personality traits remain poorly understood. To address this, we investigate how various factors, such as cultural norms and environmental stressors, encoded within LLMs, shape their personality traits, guided by the theoretical framework of social determinism. Inspired by related work on LLM interpretability, we propose a training-free approach to modify the model's behavior by extracting and steering latent features corresponding to factors within the model, thereby eliminating the need for retraining. Furthermore, we analyze the implications of these factors for model safety, focusing on their impact through the lens of personality.

Shu Yang, Shenzhe Zhu, Liang Liu, Lijie Hu, Mengdi Li, Di Wang
arXiv:2410.10863 · cs.CL, cs.AI · submitted Oct 7, 2024 · updated Feb 16, 2025
abstract · pdf · html · under review

add comment on HN