In plain words: Comparing how two neural networks represent the same data can show whether they learn alike. The usual measure fails when the representation has more dimensions than data points; an index comparing how example pairs look alike reliably matches networks trained from different random starts.
Abstract
Recent work has sought to understand the behavior of neural networks by comparing representations between layers and between different trained models. We examine methods for comparing neural network representations based on canonical correlation analysis (CCA). We show that CCA belongs to a family of statistics for measuring multivariate similarity, but that neither CCA nor any other statistic that is invariant to invertible linear transformation can measure meaningful similarities between representations of higher dimension than the number of data points. We introduce a similarity index that measures the relationship between representational similarity matrices and does not suffer from this limitation. This similarity index is equivalent to centered kernel alignment (CKA) and is also closely connected to CCA. Unlike CCA, CKA can reliably identify correspondences between representations in networks trained from different initializations.
Simon Kornblith, Mohammad Norouzi, Honglak Lee, Geoffrey Hinton
arXiv:1905.00414 · cs.LG, q-bio.NC, stat.ML · submitted May 1, 2019 · updated Jul 19, 2019
abstract · pdf · html · ICML 2019