about
CatLIP: Clip Vision Accuracy with 2.7x Faster Pre-Training on Web-Scale Data (arxiv.org)
48 points by panabee on Apr 25, 2024 | hide | past | pdf | 4 comments on HN

In plain words: Instead of comparing every image with every caption to learn matching, this approach treats each caption as a label the image must predict. It trains 2.7 times faster than the usual contrastive approach while keeping strong accuracy on tasks like detection and segmentation.

Abstract · CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data

Contrastive learning has emerged as a transformative method for learning effective visual representations through the alignment of image and text embeddings. However, pairwise similarity computation in contrastive loss between image and text pairs poses computational challenges. This paper presents a novel weakly supervised pre-training of vision models on web-scale image-text data. The proposed method reframes pre-training on image-text data as a classification task. Consequently, it eliminates the need for pairwise similarity computations in contrastive loss, achieving a remarkable $2.7\times$ acceleration in training speed compared to contrastive learning on web-scale data. Through extensive experiments spanning diverse vision tasks, including detection and segmentation, we demonstrate that the proposed method maintains high representation quality. Our source code along with pre-trained model weights and training recipes is available at \url{https://github.com/apple/corenet}.

Sachin Mehta, Maxwell Horton, Fartash Faghri, Mohammad Hossein Sekhavat, Mahyar Najibi, Mehrdad Farajtabar, Oncel Tuzel, Mohammad Rastegari
arXiv:2404.15653 · cs.CV, cs.AI, cs.CL, cs.LG · submitted Apr 24, 2024
abstract · pdf · html

add comment on HN

question: any good on-device size image embedding models?

tried https://github.com/unum-cloud/uform which i do like, especially they also support languages other than English. Any recommendations on other alternatives?

I have successfully used OpenCLIP models for embedding and similar-image search. The smallest model listed on that UForm page is 79 million parameters, so I presume that you can use other models of similar size. There are a few OpenCLIP models with 80 million or fewer parameters listed here:

https://github.com/mlfoundations/open_clip/blob/main/docs/mo...

When embeddings are quantized to int8 they still work very well for similarity (no differences in top 10 search on my test set). I haven't tried quantizing the models themselves.

TL;DR: The authors pretrain the model to classify images into Wordnet synsets[a] that appear in the caption, using a standard Cross Entropy loss. They keep the number of classes relatively small by removing any synsets that don't show up in captions at least 500 times in the dataset. It seems to work well.

My immediate question is: Why not classify among the entire hierarchy of all Wordnet synsets?

---

[a] https://wordnet.princeton.edu/

I've tried this for Wordnet hierarchical classification with a Standard Cross Entropy loss:

https://github.com/glassroom/heinsen_tree#sample-usage-with-...

It worked for me, but I had to modify the code to use all hypernym paths, giving me 147,200 classes, one per path. English only. For synsets with more than one path, I split target probability mass over their paths. For prediction, I added the predicted probs of hypernym paths ending at each synset.