In plain words: Images and text are placed in a curved space where distance grows outward, so broad ideas sit near the examples they contain, like tree branches. It kept up with the usual flat model on image classification and text-image search while making the hierarchy clear.
Abstract
Visual and linguistic concepts naturally organize themselves in a hierarchy, where a textual concept "dog" entails all images that contain dogs. Despite being intuitive, current large-scale vision and language models such as CLIP do not explicitly capture such hierarchy. We propose MERU, a contrastive model that yields hyperbolic representations of images and text. Hyperbolic spaces have suitable geometric properties to embed tree-like data, so MERU can better capture the underlying hierarchy in image-text datasets. Our results show that MERU learns a highly interpretable and structured representation space while being competitive with CLIP's performance on standard multi-modal tasks like image classification and image-text retrieval. Our code and models are available at https://www.github.com/facebookresearch/meru
Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, Ramakrishna Vedantam
arXiv:2304.09172 · cs.CV, cs.LG · submitted Apr 18, 2023 · updated Jan 18, 2024
abstract · pdf · html · ICML 2023 (v3: Add link to code in abstract)