In plain words: Swapping function names in short Python programs should not change what the code does, so the study checked whether code-writing models still solve them. They mostly fail, and bigger models grow more confident in their wrong answers instead of better.
Abstract · The Larger They Are, the Harder They Fail: Language Models do not Recognize Identifier Swaps in Python
Large Language Models (LLMs) have successfully been applied to code generation tasks, raising the question of how well these models understand programming. Typical programming languages have invariances and equivariances in their semantics that human programmers intuitively understand and exploit, such as the (near) invariance to the renaming of identifiers. We show that LLMs not only fail to properly generate correct Python code when default function names are swapped, but some of them even become more confident in their incorrect predictions as the model size increases, an instance of the recently discovered phenomenon of Inverse Scaling, which runs contrary to the commonly observed trend of increasing prediction quality with increasing model size. Our findings indicate that, despite their astonishing typical-case performance, LLMs still lack a deep, abstract understanding of the content they manipulate, making them unsuitable for tasks that statistically deviate from their training data, and that mere scaling is not enough to achieve such capability.
Antonio Valerio Miceli-Barone, Fazl Barez, Ioannis Konstas, Shay B. Cohen
arXiv:2305.15507 · cs.CL, cs.AI · submitted May 24, 2023
abstract · pdf · html · 17 pages, 5 figure, ACL 2023
What can we conclude from that kind of sequence? We can conclude that neither the original assertion, that "LLMs can't do X", nor the rebuttal, "They can if you tweak the prompt", are really telling us anything about the capabilities of LLMs to do what their users ask them to do: they are only telling us something about the capability of a user to get an LLM to do what the user wants.
In other words, we can see LLMs as random-access memories with imperfect recall. Much like a SQL query to a relational database, the right prompt can access the right piece of data, but unlike SQL nobody has a clue what "the right prompt" is in the general case.
Seen another way, any prompt a user makes to an LLM has some probability to elicit the desired response from the LLM. There is no known way to maximise that probability. Until there is, we cannot draw any conclusions about the capabilities of LLMs just by poking them and checking the results. To be very clear about it: we can't conclude either that "LLMs can't do X", nor that "LLMs can do X'. All we can conclude is "a user can do X".