In plain words: PaliGemma pairs an image-reading network with a text model and trains them broadly so the system can be adapted to many different jobs. It performed strongly across nearly 40 tasks, including satellite-image reading and outlining objects in pictures.
Abstract
PaliGemma is an open Vision-Language Model (VLM) that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model. It is trained to be a versatile and broadly knowledgeable base model that is effective to transfer. It achieves strong performance on a wide variety of open-world tasks. We evaluate PaliGemma on almost 40 diverse tasks including standard VLM benchmarks, but also more specialized tasks such as remote-sensing and segmentation.
Lucas Beyer, Andreas Steiner, André Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, et al.
arXiv:2407.07726 · cs.CV, cs.AI, cs.CL, cs.LG · submitted Jul 10, 2024 · updated Oct 10, 2024
abstract · pdf · html · v2 adds Appendix H and I and a few citations