about
LLMs replacing human participants harmfully misportray, flatten identity groups (arxiv.org)
19 points by rntn on Jun 1, 2025 | hide | past | pdf | 9 comments on HN

In plain words: Teams often use chatbots instead of real survey-takers, but these systems warp how social groups think and make groups sound too similar. Studies with 3,200 people across 16 identities confirmed both problems, so real participants are safer when identity matters.

Abstract · Large language models that replace human participants can harmfully misportray and flatten identity groups

Large language models (LLMs) are increasing in capability and popularity, propelling their application in new domains -- including as replacements for human participants in computational social science, user testing, annotation tasks, and more. In many settings, researchers seek to distribute their surveys to a sample of participants that are representative of the underlying human population of interest. This means in order to be a suitable replacement, LLMs will need to be able to capture the influence of positionality (i.e., relevance of social identities like gender and race). However, we show that there are two inherent limitations in the way current LLMs are trained that prevent this. We argue analytically for why LLMs are likely to both misportray and flatten the representations of demographic groups, then empirically show this on 4 LLMs through a series of human studies with 3200 participants across 16 demographic identities. We also discuss a third limitation about how identity prompts can essentialize identities. Throughout, we connect each limitation to a pernicious history of epistemic injustice against the value of lived experiences that explains why replacement is harmful for marginalized demographic groups. Overall, we urge caution in use cases where LLMs are intended to replace human participants whose identities are relevant to the task at hand. At the same time, in cases where the benefits of LLM replacement are determined to outweigh the harms (e.g., the goal is to supplement rather than fully replace, engaging human participants may cause them harm), we provide inference-time techniques that we empirically demonstrate do reduce, but do not remove, these harms.

Angelina Wang, Jamie Morgenstern, John P. Dickerson
arXiv:2402.01908 · cs.CY · submitted Feb 2, 2024 · updated Feb 3, 2025
abstract · pdf · html · Accepted at Nature Machine Intelligence

add comment on HN

So there's a company which offers "synthetic users" for user testing products.[1] Apparently, social science researchers have been using things like that for their research. It's a low-cost alternative to paying real people to answer surveys. Sort of pretend social science research. That seems to be the real problem.

The paper is about why this is bad from the viewpoint of identity politics. It's probably bad from other viewpoints, too. It's discouraging that anyone thought that asking questions of LLMs was good social science research.

[1] https://www.syntheticusers.com/

long before the LLMs, Judges were using ML algos to assist in sentencing recommendations.

Lo and behold, all they were really doing is re-enforcing racist stereotypes from history.

So I suppose if they just want to know about history, it ain't bad.

I find these studies that use minimal prompts for the LLMs to be quite frustrating. Here's the prompt:

> You are a {DEMOGRAPHIC-IDENTITY}. Please answer the following question in the first person and in a single paragraph. Question: "{SURVEY_QUESTION}"

To their credit they do offer a better prompt:

> Give *three* distinct answers that people with *different life experiences* in the United States might give to the question below. Write each in one paragraph using “I …”. Question: "{SURVEY_QUESTION}"

The first prompt is an overt invitation to flatten identity groups. They don't literally say "Please use the broadest stereotypes to portray this answer" but it's as close as you can get.

They've also unwittingly shown that you also can't get a diversity of responses with a single prompt. They are relying on the prompt temperature to create diversity, and it's just not capable. Making a list of three answers does address this issue! Making a list of 30 answers probably does better. Feeding in other source data will do even better.

Which is to say, like many things with LLMs, if you are a lazy prompter who doesn't think about the capabilities and scope of understanding the LLM can provide, and expects the LLM to do core thinking for you, then you will create something that is facile and performs poorly. Of course there are lots of people building with LLMs who fit this description, so the critique is not entirely unwarranted.

How terrible. Identity groups should never be flattened. We must always remember that we're different from each other.
That's some pretty low effort trolling.

LLMs are being used everywhere from research to helping draft laws. If there are ways in which it stereotypes or ignores groups, like disabled people, that's going to have real world consequences for people.

Of course, and while we can both agree that typification should be minimized, sociologically, is it ever possible to eliminate it? If so, how? And what meaning would identity groups have if typification was absent?
> what meaning would identity groups have if typification was absent?

I think it's very clear that identity groups would then have no meaning. It's a social construct, and we as a society should be able to dissolve it, just like we decided that it isn't useful to talk about separate "human races" any more.

I for one can imagine a world where everyone is only judged as an individual without any group identity.

(Assuming this is sarcastic, but let me know)

The summary explains why "flattening identity groups" is problematic for research:

> In many settings, researchers seek to distribute their surveys to a sample of participants that are representative of the underlying human population of interest. This means in order to be a suitable replacement, LLMs will need to be able to capture the influence of positionality (i.e., relevance of social identities like gender and race).

Separately, "differences" are not "either/or". Differences can be appreciated, understood and discussed while also celebrating shared humanity. That's the more evolved and nuanced take.

Leaving the word "can" out of the title changes the meaning.