about
Emerging Properties in Unified Multimodal Pretraining (arxiv.org)
1 point by buildbot on May 21, 2025 | hide | past | pdf | 1 comment on HN

In plain words: A single open model trained on trillions of mixed text, image, video, and web tokens both understands and creates images and video. It beats other open unified models at both jobs and gains new skills like free-form image editing and predicting future video frames.

Abstract

Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. In this work, we introduce BAGEL, an open-source foundational model that natively supports multimodal understanding and generation. BAGEL is a unified, decoder-only model pretrained on trillions of tokens curated from large-scale interleaved text, image, video, and web data. When scaled with such diverse multimodal interleaved data, BAGEL exhibits emerging capabilities in complex multimodal reasoning. As a result, it significantly outperforms open-source unified models in both multimodal generation and understanding across standard benchmarks, while exhibiting advanced multimodal reasoning abilities such as free-form image manipulation, future frame prediction, 3D manipulation, and world navigation. In the hope of facilitating further opportunities for multimodal research, we share the key findings, pretraining details, data creation protocal, and release our code and checkpoints to the community. The project page is at https://bagel-ai.org/

Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan
arXiv:2505.14683 · cs.CV · submitted May 20, 2025 · updated Jul 27, 2025
abstract · pdf · html · 37 pages, 17 figures

add comment on HN

New Multimodal in input & output 14B param LLM from Bytedance. Supports chain of thought reasoning and image generation.