In plain words: They tested how language models handle gender-neutral pronouns in Danish, English, and Swedish by measuring how confidently the models predict text and answer questions. The models stumbled on these pronouns, scoring worse than with older gendered ones, even though people read them easily.
Abstract · How Conservative are Language Models? Adapting to the Introduction of Gender-Neutral Pronouns
Gender-neutral pronouns have recently been introduced in many languages to a) include non-binary people and b) as a generic singular. Recent results from psycholinguistics suggest that gender-neutral pronouns (in Swedish) are not associated with human processing difficulties. This, we show, is in sharp contrast with automated processing. We show that gender-neutral pronouns in Danish, English, and Swedish are associated with higher perplexity, more dispersed attention patterns, and worse downstream performance. We argue that such conservativity in language models may limit widespread adoption of gender-neutral pronouns and must therefore be resolved.
Stephanie Brandl, Ruixiang Cui, Anders Søgaard
arXiv:2204.10281 · cs.CL · submitted Apr 11, 2022 · updated May 3, 2022
abstract · pdf · html · To appear at NAACL 2022
If pronoun X were as commonly used as "he" or "she" than a language model would learn it as well as "he" or "she". If X is used 0.1% as often then the language model is going to have trouble.
You could make a synthetic data set where "he", "she" and X occur all about 1/3 of time and the model would handle X pretty well in certain ways, but would make other mistakes because it thinks X occurs too often.
It's a basic problem of language models that they conflate syntax and semantics and thus get the meaning of things wrong. They come to conclusions like "Tyrone is a thug", "Fat people are transgender," etc. (When do "they" get to put A on their passport for Adipose?)