about
Why Are All LLMs Obsessed with Japanese Culture? (arxiv.org)
4 points by geox 149 days ago | hide | past | pdf | 2 comments on HN

In plain words: They asked chatbots open-ended culture questions in 24 languages, each time requesting a sample place, to see which countries they favor. The models leaned toward Japan, and this preference appeared only after the step where models learn to follow instructions, not during initial training.

Abstract · Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs

LLMs have limitations when it comes to cultural coverage and competence, and in some cases, show specific cultural biases. Although prior studies have examined the cultural capabilities of LLMs, none have specifically investigated their regional preferences in generic culture-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ), with questions available in 24 languages. We evaluate LLMs by prompting them to answer questions from CROQ and provide a sample location. The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan in their answers. Moreover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs. Low-resource languages, on the other hand, show more inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training. Dataset available at https://huggingface.co/datasets/HiTZ/CROQ

Joseba Fernandez de Landa, Carla Perez-Almendros, Jose Camacho-Collados
arXiv:2604.21751 · cs.CL, cs.AI, cs.CY · submitted Apr 23, 2026 · updated Aug 28, 2026
abstract · pdf · html

add comment on HN

So they use LLM to evaluate LLMs: with LLM writing the questions, another LLM writing the country-specific answers, and yet another LLM getting the country from an answer. The only manual steps seem to be "manually reviewed [questions] to remove repetitions or accidental location references."

This seems like a pretty lazy methodology, as if there are LLM-specific country biases, they could be introduced at any stage of the process.

As a Japanese editor, this research feels like it has finally put words to the "discomfort" I’ve been sensing.

In Japanese, the most meaningful parts of a text often reside in the "Ma" (space) or in the unspoken context. However, because the text AI presents as "correct" seems to have passed through a Western logical filter, it feels as though cultural nuances are being treated as "logical flaws" or "ambiguities."

If this continues, the internet may become flooded with uninteresting writing that fails to move anyone’s heart.