about
The Persistence of Neural Collapse Despite Low-Rank Bias (arxiv.org)
1 point by Wheatman on Nov 4, 2024 | hide | past | pdf | 1 comment on HN

In plain words: Trained networks often squeeze their features into a neat, symmetric pattern called neural collapse, though theory says it isn't the best fit. The study finds this pattern is more common across a network's possible settings than other arrangements, explaining why it keeps appearing.

Abstract

Neural collapse (NC) and its multi-layer variant, deep neural collapse (DNC), describe a structured geometry that occurs in the features and weights of trained deep networks. Recent theoretical work by Sukenik et al. using a deep unconstrained feature model (UFM) suggests that DNC is suboptimal under mean squared error (MSE) loss. They heuristically argue that this is due to low-rank bias induced by L2 regularization. In this work, we extend this result to deep UFMs trained with cross-entropy loss, showing that high-rank structures, including DNC, are not generally optimal. We characterize the associated low-rank bias, proving a fixed bound on the number of non-negligible singular values at global minima as network depth increases. We further analyze the loss surface, demonstrating that DNC is more prevalent in the landscape than other critical configurations, which we argue explains its frequent empirical appearance. Our results are validated through experiments in deep UFMs and deep neural networks.

Connall Garrod, Jonathan P. Keating
arXiv:2410.23169 · cs.LG · submitted Oct 30, 2024 · updated Oct 5, 2025
abstract · pdf · html · 48 pages, 19 figures. To appear in NeurIPS 2025. Slightly reformatted for readability

add comment on HN

Note: This isn't to be confused with model collapse (the rapid deterioration of models trained on their own data), this paper talks about neural collapse (Read here:https://medium.com/ai-assimilating-intelligence/what-is-neur...)