In plain words: Children, adults, and AI chatbots solved letter-sequence puzzles like "a b : a c :: j k : ?" in Latin, then in Greek letters and plain symbols. People carried the rule over easily to the new alphabets, but the AI systems mostly failed.
Abstract · Can Large Language Models generalize analogy solving like children can?
In people, the ability to solve analogies such as "body : feet :: table : ?" emerges in childhood, and appears to transfer easily to other domains, such as the visual domain "( : ) :: < : ?". Recent research shows that large language models (LLMs) can solve various forms of analogies. However, can LLMs generalize analogy solving to new domains like people can? To investigate this, we had children, adults, and LLMs solve a series of letter-string analogies (e.g., a b : a c :: j k : ?) in the Latin alphabet, in a near transfer domain (Greek alphabet), and a far transfer domain (list of symbols). Children and adults easily generalized their knowledge to unfamiliar domains, whereas LLMs did not. This key difference between human and AI performance is evidence that these LLMs still struggle with robust human-like analogical transfer.
Claire E. Stevenson, Alexandra Pafford, Han L. J. van der Maas, Melanie Mitchell
arXiv:2411.02348 · cs.AI, cs.CL, cs.HC · submitted Nov 4, 2024 · updated Oct 6, 2025
abstract · pdf · html · Accepted to Transactions of the Association for Computational Linguistics (TACL)
In contrast, in children, familiarity with letters or symbols does not seem to influence how well they solve letter-string analogies. As such, our results add to the accumulating evidence that questions whether reasoning actually occurs in these LLMs
Not terribly surprising I would think, but a nice study to have in case someone has had to much Kool-Aid.