about
I-Con: A Unifying Framework for Representation Learning (arxiv.org)
1 point by aerophilic on Jun 5, 2025 | hide | past | pdf | 1 comment on HN

In plain words: A single equation shows that many training losses, from clustering to contrastive learning, all do one thing: shrink the gap between what a label says and what the model learns. Losses built from it classified unlabeled ImageNet images 8% better than the previous best.

Abstract

As the field of representation learning grows, there has been a proliferation of different loss functions to solve different classes of problems. We introduce a single information-theoretic equation that generalizes a large collection of modern loss functions in machine learning. In particular, we introduce a framework that shows that several broad classes of machine learning methods are precisely minimizing an integrated KL divergence between two conditional distributions: the supervisory and learned representations. This viewpoint exposes a hidden information geometry underlying clustering, spectral methods, dimensionality reduction, contrastive learning, and supervised learning. This framework enables the development of new loss functions by combining successful techniques from across the literature. We not only present a wide array of proofs, connecting over 23 different approaches, but we also leverage these theoretical results to create state-of-the-art unsupervised image classifiers that achieve a +8% improvement over the prior state-of-the-art on unsupervised classification on ImageNet-1K. We also demonstrate that I-Con can be used to derive principled debiasing methods which improve contrastive representation learners.

Shaden Alshammari, John Hershey, Axel Feldmann, William T. Freeman, Mark Hamilton
arXiv:2504.16929 · cs.LG, cs.AI, cs.CV, cs.IT · submitted Apr 23, 2025
abstract · pdf · html · ICLR 2025; website: https://aka.ms/i-con . Proceedings of the Thirteenth International Conference on Learning Representations (ICLR 2025)

add comment on HN

This seems to be a way to identify and create a 'periodic table' of different machine learning approaches... importantly identifying 'gaps' between approaches which may yield new ways of doing various learning techniques.