In plain words: As AI models grow larger and see more data, their ways of organizing information become more alike across image and text systems. The bigger they get, the more similarly they judge which data points are close, hinting at one shared picture of reality.
Abstract · The Platonic Representation Hypothesis
We argue that representations in AI models, particularly deep networks, are converging. First, we survey many examples of convergence in the literature: over time and across multiple domains, the ways by which different neural networks represent data are becoming more aligned. Next, we demonstrate convergence across data modalities: as vision models and language models get larger, they measure distance between datapoints in a more and more alike way. We hypothesize that this convergence is driving toward a shared statistical model of reality, akin to Plato's concept of an ideal reality. We term such a representation the platonic representation and discuss several possible selective pressures toward it. Finally, we discuss the implications of these trends, their limitations, and counterexamples to our analysis.
Minyoung Huh, Brian Cheung, Tongzhou Wang, Phillip Isola
arXiv:2405.07987 · cs.LG, cs.AI, cs.CV, cs.NE · submitted May 13, 2024 · updated Jul 25, 2024
abstract · pdf · html · Equal contributions. Project: https://phillipi.github.io/prh/ Code: https://github.com/minyoungg/platonic-rep
If we had these representations converge in a real computing neural network, maybe then we could argue that they were Plato's ideals a step above the corrupted quantization of physical forms. But as they are right now, the representations are an even more reduced representations of the physical originals.
Still very neat that it's happening though.