about
Dialect prejudice predicts AI decisions about people's character (arxiv.org)
4 points by saassiopeia on Mar 5, 2024 | hide | past | pdf | 2 comments on HN

In plain words: Tests showed AI chatbots carry hidden prejudice against African American English speakers, rating them more negatively than any human stereotype ever recorded. Given only how someone talks, they steered speakers toward lower jobs, guilty verdicts, and death sentences; bias training did not fix it.

Abstract · Dialect prejudice predicts AI decisions about people's character, employability, and criminality

Hundreds of millions of people now interact with language models, with uses ranging from serving as a writing aid to informing hiring decisions. Yet these language models are known to perpetuate systematic racial prejudices, making their judgments biased in problematic ways about groups like African Americans. While prior research has focused on overt racism in language models, social scientists have argued that racism with a more subtle character has developed over time. It is unknown whether this covert racism manifests in language models. Here, we demonstrate that language models embody covert racism in the form of dialect prejudice: we extend research showing that Americans hold raciolinguistic stereotypes about speakers of African American English and find that language models have the same prejudice, exhibiting covert stereotypes that are more negative than any human stereotypes about African Americans ever experimentally recorded, although closest to the ones from before the civil rights movement. By contrast, the language models' overt stereotypes about African Americans are much more positive. We demonstrate that dialect prejudice has the potential for harmful consequences by asking language models to make hypothetical decisions about people, based only on how they speak. Language models are more likely to suggest that speakers of African American English be assigned less prestigious jobs, be convicted of crimes, and be sentenced to death. Finally, we show that existing methods for alleviating racial bias in language models such as human feedback training do not mitigate the dialect prejudice, but can exacerbate the discrepancy between covert and overt stereotypes, by teaching language models to superficially conceal the racism that they maintain on a deeper level. Our findings have far-reaching implications for the fair and safe employment of language technology.

Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, Sharese King
arXiv:2403.00742 · cs.CL, cs.AI, cs.CY · submitted Mar 1, 2024
abstract · pdf · html

add comment on HN
Also discussed: Mar 2024 (1 point, 1 comment)

A very important result from this study is that GPT-4 is in many ways worse than GPT-3.5, which was worse than GPT-3. More RLHF makes LLMs hide the overt racism better and leads to a false sense of security ("thank God GPT finally stopped using n-bombs") but they are learning more covert racism from their giant pile of badly curated training data, and this covert bias is not being adequately addressed in RLHF.

Any company using LLMs for hiring decisions needs to be investigated by the feds.

How... unsurprising...