about
Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories (arxiv.org)
4 points by danielrmay 117 days ago | hide | past | pdf | 1 comment on HN

In plain words: Across 20,000 stories from four AI models, 11 words — names like Elias, settings like lighthouses — appeared in 88.3% of them. These words appear in the small sets of human-picked examples used to train models, so a handful of examples can steer what models write.

Abstract

LLM-generated stories are a popular use case, but they show very low variability. We sample 20,000 total stories from four current models using five prompts. We find that 11 words occur in 88.3% of generated stories, with little difference between models. These words include names (Elias, Mara, Elara), settings (lighthouses), and professions (clockmaker, librarian). These tokens do not often occur in published literature nor pre-training data, but they are found in preference data that is likely to have been used by all current models. Surprisingly, these "lighthouse" stories are infrequent when compared with the average post-training story, much of which contains references to copyrighted characters or adult content. This result demonstrates the potentially disproportionate impact of small datasets combined with powerful alignment algorithms.

Sil Hamilton, David Mimno
arXiv:2605.26492 · cs.CL, cs.AI, cs.LG · submitted May 26, 2026
abstract · pdf · html

add comment on HN

About a month ago I wrote a blog post that spoke to the rise of poor quality content as a result of agentic workflows, but it primarily focused on ethics: ending with presenting how a synthetic author named "Elias Thorne" is now rising in _alternative Cancer treatment_ book categories on Amazon, and questioning how this will affect the web over time. That post is here if you're interested: https://danielmay.co.uk/posts/cheap-agents-alumni-shirts-and...

A few weeks after, a doctoral research team from Cornell published this paper to arXiv that digs into more of the same phenomenon scientifically (n=20k vs my 8!) and examines the training data to try to understand cause. I'm interested to continue to see what the second-order effects of this phenomena will be.