about
Self-Supervised Learning from Images with JEPA (2023) (arxiv.org)
40 points by Brysonbw on Mar 29, 2025 | hide | past | pdf | 10 comments on HN

In plain words: It learns image features by hiding parts of a picture and guessing the hidden parts' abstract features from one visible region, rather than hand-tuned crops and color tweaks. A huge version trained in under 72 hours and did well counting objects and judging depth.

Abstract · Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

This paper demonstrates an approach for learning highly semantic image representations without relying on hand-crafted data-augmentations. We introduce the Image-based Joint-Embedding Predictive Architecture (I-JEPA), a non-generative approach for self-supervised learning from images. The idea behind I-JEPA is simple: from a single context block, predict the representations of various target blocks in the same image. A core design choice to guide I-JEPA towards producing semantic representations is the masking strategy; specifically, it is crucial to (a) sample target blocks with sufficiently large scale (semantic), and to (b) use a sufficiently informative (spatially distributed) context block. Empirically, when combined with Vision Transformers, we find I-JEPA to be highly scalable. For instance, we train a ViT-Huge/14 on ImageNet using 16 A100 GPUs in under 72 hours to achieve strong downstream performance across a wide range of tasks, from linear classification to object counting and depth prediction.

Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, Nicolas Ballas
arXiv:2301.08243 · cs.CV, cs.AI, cs.LG, eess.IV · submitted Jan 19, 2023 · updated Apr 13, 2023
abstract · pdf · html · 2023 IEEE/CVF International Conference on Computer Vision

add comment on HN
Also discussed: Jun 2024 (1 point, 0 comments)

It’s not new and only superior in a very narrow set of categories.
As a computer vision guy I'm sad JEPA didn't end up more effective. Makes perfect sense conceptually, would have easily transferred to video, but other self-supervised methods just seem to beat it!
Yeah! JEPA seems awesome. Do you mind sharing what other self-supervised methods work better than JEPA?
Needs a (2023) tag. But definitely the release of ARC2 and image outputs from 4o got me thinking about the JEPA family too.

I don't know if it's right (and I'm sure JEPA has lots of performance issues) but seems good to have a fully latent space representation, ideally across all modalities, so that the concept "an apple a day keeps the doctor away" becoming image/audio/text is a choice of decoder rather than dedicated token ranges being chosen even before the actual creation process in the model begins.

GPTs are in the “exploit” phase of the “explore-exploit” trade-off.

JEPA is still in the explore phase, it’s good to read the paper and have an understanding of the architecture to gain an alternative perspective.

Not new, not notable right now, not sure why it's getting upvoted (just kidding, it's because people see YLC and upvote based on names)
Even average papers can have nice overview of the problem and references.
I don't care for names, i just thought it was an interesting read.
JEPA is presumably superior to Transformers. Can any expert enlighten us on the implications of this paper?
Transformers are usually part of JEPA architectures. In I-JEPA's case, there is a ViT that is used in the context encoding phase.