In plain words: They tested AI chatbots on text with the letters inside every word jumbled, asking them to rebuild the sentences and answer questions about them. GPT-4 almost perfectly restored the originals, cutting the remaining letter errors by 95%, while other models and even people often failed.
Abstract · Unnatural Error Correction: GPT-4 Can Almost Perfectly Handle Unnatural Scrambled Text
While Large Language Models (LLMs) have achieved remarkable performance in many tasks, much about their inner workings remains unclear. In this study, we present novel experimental insights into the resilience of LLMs, particularly GPT-4, when subjected to extensive character-level permutations. To investigate this, we first propose the Scrambled Bench, a suite designed to measure the capacity of LLMs to handle scrambled input, in terms of both recovering scrambled sentences and answering questions given scrambled context. The experimental results indicate that most powerful LLMs demonstrate the capability akin to typoglycemia, a phenomenon where humans can understand the meaning of words even when the letters within those words are scrambled, as long as the first and last letters remain in place. More surprisingly, we found that only GPT-4 nearly flawlessly processes inputs with unnatural errors, even under the extreme condition, a task that poses significant challenges for other LLMs and often even for humans. Specifically, GPT-4 can almost perfectly reconstruct the original sentences from scrambled ones, decreasing the edit distance by 95%, even when all letters within each word are entirely scrambled. It is counter-intuitive that LLMs can exhibit such resilience despite severe disruption to input tokenization caused by scrambled text.
Qi Cao, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
arXiv:2311.18805 · cs.CL, cs.AI · submitted Nov 30, 2023
abstract · pdf · html · EMNLP 2023 (with an additional analysis section in appendix)
This was interesting because word segmentation is a difficult problem that is usually thought to require something like dynamic programming[1][2] to get right. It's a little surprising that GPT-4 can handle this, because it has no capability to search different alternatives to backtrack if it makes a mistake, but apparently it's stronger understanding of language means that it doesn't really need to.
It's also surprising that tokenization doesn't appear to interfere with its ability to these tasks, because it seems like it would make things a lot harder. According to the openAI tokenizer[3], GPT-4 sees the following tokens in the above text:
Except for "UNDER", "SEA", and "OF", almost all of those token breaks are not at natural word boundaries. The same is true for the scrambled text examples in the original article. So GPT-4 must actually be taking those tokens apart into individual letters and gluing them back together into completely new tokens somewhere inside it's many layers of transformers.[1]: https://web.cs.wpi.edu/~cs2223/b05/HW/HW6/SolutionsHW6/
[2]: https://pypi.org/project/wordsegmentation/
[3]: https://platform.openai.com/tokenizer