In plain words: It squeezes data into a smaller map while keeping its clusters, loops, and holes intact, training an autoencoder to minimize a score measuring how much those shapes change under compression. It beat the best competing methods at keeping the data's shape on three checks.
Abstract
We propose a method for learning topology-preserving data representations (dimensionality reduction). The method aims to provide topological similarity between the data manifold and its latent representation via enforcing the similarity in topological features (clusters, loops, 2D voids, etc.) and their localization. The core of the method is the minimization of the Representation Topology Divergence (RTD) between original high-dimensional data and low-dimensional representation in latent space. RTD minimization provides closeness in topological features with strong theoretical guarantees. We develop a scheme for RTD differentiation and apply it as a loss term for the autoencoder. The proposed method "RTD-AE" better preserves the global structure and topology of the data manifold than state-of-the-art competitors as measured by linear correlation, triplet distance ranking accuracy, and Wasserstein distance between persistence barcodes.
Ilya Trofimov, Daniil Cherniavskii, Eduard Tulchinskii, Nikita Balabin, Evgeny Burnaev, Serguei Barannikov
arXiv:2302.00136 · cs.LG, cs.AI, math.AT, math.KT, math.SG · submitted Jan 31, 2023 · updated Feb 15, 2023
abstract · pdf · html