about
Questioning Representational Optimism in Deep Learning (arxiv.org)
1 point by vatsachak 323 days ago | hide | past | pdf | 1 comment on HN

In plain words: Networks evolved by open-ended search were compared with ones trained the usual way to make an image, picturing neurons to see how the output forms. Both made the same image, but usual training left the insides jumbled and tangled; the evolved ones stayed clear.

Abstract · Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis

Much of the excitement in modern AI is driven by the observation that scaling up existing systems leads to better performance. But does better performance necessarily imply better internal representations? While the representational optimist assumes it must, this position paper challenges that view. We compare neural networks evolved through an open-ended search process to networks trained via conventional stochastic gradient descent (SGD) on the simple task of generating a single image. This minimal setup offers a unique advantage: each hidden neuron's full functional behavior can be easily visualized as an image, thus revealing how the network's output behavior is internally constructed neuron by neuron. The result is striking: while both networks produce the same output behavior, their internal representations differ dramatically. The SGD-trained networks exhibit a form of disorganization that we term fractured entangled representation (FER). Interestingly, the evolved networks largely lack FER, even approaching a unified factored representation (UFR). In large models, FER may be degrading core model capacities like generalization, creativity, and (continual) learning. Therefore, understanding and mitigating FER could be critical to the future of representation learning.

Akarsh Kumar, Jeff Clune, Joel Lehman, Kenneth O. Stanley
arXiv:2505.11581 · cs.CV, cs.LG, cs.NE · submitted May 16, 2025
abstract · pdf · html · 43 pages, 25 figures

add comment on HN
Also discussed: Jun 2025 (1 point, 3 comments) · May 2025 (1 point, 0 comments)

The above paper uses two methods to construct the same image, one through SGD with the loss against pixels and the other through an "evolutionary" method. Although both methods construct the image, the evolutionary method stores better representations of how to construct the image in its weights.