about
"Yes I Would Recommend Calling the Police":Norm Inconsistency in LLM Decisions (arxiv.org)
1 point by belter on May 24, 2024 | hide | past | pdf | discuss on HN

In plain words: Three AI chatbots watched home-surveillance clips and decided whether to call the police, while the study checked their answers against the activity shown and the people and neighborhood. Their advice often ignored actual crimes and shifted with the neighborhood's racial makeup.

Abstract · As an AI Language Model, "Yes I Would Recommend Calling the Police": Norm Inconsistency in LLM Decision-Making

We investigate the phenomenon of norm inconsistency: where LLMs apply different norms in similar situations. Specifically, we focus on the high-risk application of deciding whether to call the police in Amazon Ring home surveillance videos. We evaluate the decisions of three state-of-the-art LLMs -- GPT-4, Gemini 1.0, and Claude 3 Sonnet -- in relation to the activities portrayed in the videos, the subjects' skin-tone and gender, and the characteristics of the neighborhoods where the videos were recorded. Our analysis reveals significant norm inconsistencies: (1) a discordance between the recommendation to call the police and the actual presence of criminal activity, and (2) biases influenced by the racial demographics of the neighborhoods. These results highlight the arbitrariness of model decisions in the surveillance context and the limitations of current bias detection and mitigation strategies in normative decision-making.

Shomik Jain, D Calacci, Ashia Wilson
arXiv:2405.14812 · cs.CY · submitted May 23, 2024 · updated Aug 17, 2024
abstract · pdf · html · To appear in the proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2024)

add comment on HN