about
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics (arxiv.org)
5 points by XTXinverseXTY 325 days ago | hide | past | pdf | 2 comments on HN

In plain words: A self-supervised recipe shapes a model's internal representations to spread evenly in every direction, which theory says best lowers later prediction error, while dropping tricks like teacher networks and gradient blocking. With a frozen backbone it scored 79% on ImageNet and stayed stable across architectures.

Abstract

Learning manipulable representations of the world and its dynamics is central to AI. Joint-Embedding Predictive Architectures (JEPAs) offer a promising blueprint, but lack of practical guidance and theory has led to ad-hoc R&D. We present a comprehensive theory of JEPAs and instantiate it in {\bf LeJEPA}, a lean, scalable, and theoretically grounded training objective. First, we identify the isotropic Gaussian as the optimal distribution that JEPAs' embeddings should follow to minimize downstream prediction risk. Second, we introduce a novel objective--{\bf Sketched Isotropic Gaussian Regularization} (SIGReg)--to constrain embeddings to reach that ideal distribution. Combining the JEPA predictive loss with SIGReg yields LeJEPA with numerous theoretical and practical benefits: (i) single trade-off hyperparameter, (ii) linear time and memory complexity, (iii) stability across hyper-parameters, architectures (ResNets, ViTs, ConvNets) and domains, (iv) heuristics-free, e.g., no stop-gradient, no teacher-student, no hyper-parameter schedulers, and (v) distributed training-friendly implementation requiring only $\approx$50 lines of code. Our empirical validation covers 10+ datasets, 60+ architectures, all with varying scales and domains. As an example, using imagenet-1k for pretraining and linear evaluation with frozen backbone, LeJEPA reaches 79\% with a ViT-H/14. We hope that the simplicity and theory-friendly ecosystem offered by LeJEPA will reestablish self-supervised pre-training as a core pillar of AI research (\href{https://github.com/rbalestr-lab/lejepa}{GitHub repo}).

Randall Balestriero, Yann LeCun
arXiv:2511.08544 · cs.LG, cs.AI, cs.CV, stat.ML · submitted Nov 11, 2025 · updated Nov 14, 2025
abstract · pdf · html

add comment on HN
Also discussed: Nov 2025 (68 points, 18 comments) · Nov 2025 (2 points, 0 comments)

A breath of fresh air from the over-concentration on LLM research. Something different for once.
Seems like this signals Yann Lecun's direction now that he's leaving Meta

The EMA teacher model still seems like black magic to me (present in DINO and JEPA series but gone now)