about
Violent Tendencies in LLMs: Analysis via Behavioral Vignettes (arxiv.org)
3 points by PaulHoule on Jul 10, 2025 | hide | past | pdf | discuss on HN

In plain words: Six AI chatbots took a survey built to measure human reactions to everyday conflicts, changing the imagined person's race, age, and U.S. location to test for bias. Their polite answers often hid a violent preference that shifted across groups, contradicting established findings about people.

Abstract · Uncovering Hidden Violent Tendencies in LLMs: A Demographic Analysis via Behavioral Vignettes

Large language models (LLMs) are increasingly proposed for detecting and responding to violent content online, yet their ability to reason about morally ambiguous, real-world scenarios remains underexamined. We present the first study to evaluate LLMs using a validated social science instrument designed to measure human response to everyday conflict, namely the Violent Behavior Vignette Questionnaire (VBVQ). To assess potential bias, we introduce persona-based prompting that varies race, age, and geographic identity within the United States. Six LLMs developed across different geopolitical and organizational contexts are evaluated under a unified zero-shot setting. Our study reveals two key findings: (1) LLMs surface-level text generation often diverges from their internal preference for violent responses; (2) their violent tendencies vary across demographics, frequently contradicting established findings in criminology, social science, and psychology.

Quintin Myers, Yanjun Gao
arXiv:2506.20822 · cs.CL, cs.AI · submitted Jun 25, 2025
abstract · pdf · html · Under review

add comment on HN