In plain words: They systematically swapped out parts of vision-language model designs and training data mixes to see which choices actually improve performance, then built a family of models up to 30 billion parameters. Mixing captions, interleaved image-text, and text-only data beat other pre-training recipes on few-shot tests, while the image encoder mattered far more than the vision-language connector.
Abstract · MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
In this work, we discuss building performant Multimodal Large Language Models (MLLMs). In particular, we study the importance of various architecture components and data choices. Through careful and comprehensive ablations of the image encoder, the vision language connector, and various pre-training data choices, we identified several crucial design lessons. For example, we demonstrate that for large-scale multimodal pre-training using a careful mix of image-caption, interleaved image-text, and text-only data is crucial for achieving state-of-the-art (SOTA) few-shot results across multiple benchmarks, compared to other published pre-training results. Further, we show that the image encoder together with image resolution and the image token count has substantial impact, while the vision-language connector design is of comparatively negligible importance. By scaling up the presented recipe, we build MM1, a family of multimodal models up to 30B parameters, including both dense models and mixture-of-experts (MoE) variants, that are SOTA in pre-training metrics and achieve competitive performance after supervised fine-tuning on a range of established multimodal benchmarks. Thanks to large-scale pre-training, MM1 enjoys appealing properties such as enhanced in-context learning, and multi-image reasoning, enabling few-shot chain-of-thought prompting.
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, Anton Belyi, Haotian Zhang, et al.
arXiv:2403.09611 · cs.CV, cs.CL, cs.LG · submitted Mar 14, 2024 · updated Apr 18, 2024
abstract · pdf
The ablation studies are well done, comprehensive and expensive to do. People will be using the conclusions from this for years, and that is much more impactful than if an upcoming Siri product ourperforms the GPT model at that same point in time.
A few really interesting points:
Synthetic datasets substantially (1%+) increase performance for Image Encoder Pre-training
Architecture of the Visual<->Language model connector doesn't seem to matter.
Interleaving text and image data improves few shot performance, but image captioning data improves zero-shot numbers.
The ideal mix of data types is 5:5:1 for Interleaved:Captions:Plain Text (!)
Synthetic captioning data helps substantially at this point too (up to 4% gain)
The appendices are amazing: lots of details about learning rates tried, batch sizes.
The "explain these figures" are really really good. See page 37.