In plain words: A model learns from pairs of fashion photos and their descriptions, so it understands general fashion ideas instead of one narrow task. It then handles product search, category sorting, and spotting items in photos, where the usual approach trains a separate model for each.
Abstract · Contrastive language and vision learning of general fashion concepts
The steady rise of online shopping goes hand in hand with the development of increasingly complex ML and NLP models. While most use cases are cast as specialized supervised learning problems, we argue that practitioners would greatly benefit from more transferable representations of products. In this work, we build on recent developments in contrastive learning to train FashionCLIP, a CLIP-like model for the fashion industry. We showcase its capabilities for retrieval, classification and grounding, and release our model and code to the community.
Patrick John Chia, Giuseppe Attanasio, Federico Bianchi, Silvia Terragni, Ana Rita Magalhães, Diogo Goncalves, Ciro Greco, Jacopo Tagliabue
arXiv:2204.03972 · cs.IR, cs.CL · submitted Apr 8, 2022 · updated Apr 18, 2023
abstract · pdf · html · Latest version available at https://www.nature.com/articles/s41598-022-23052-9; model available at https://huggingface.co/patrickjohncyh/fashion-clip