about
Janus-Pro: Multimodal Understanding and Generation with Data and Model Scaling (arxiv.org)
1 point by belter on Jan 30, 2025 | hide | past | pdf | discuss on HN

In plain words: This system both understands images and draws new ones from text descriptions, using a better training recipe, more training data, and a larger size. Compared with the earlier version, it handles both jobs better and produces steadier images.

Abstract · Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

In this work, we introduce Janus-Pro, an advanced version of the previous work Janus. Specifically, Janus-Pro incorporates (1) an optimized training strategy, (2) expanded training data, and (3) scaling to larger model size. With these improvements, Janus-Pro achieves significant advancements in both multimodal understanding and text-to-image instruction-following capabilities, while also enhancing the stability of text-to-image generation. We hope this work will inspire further exploration in the field. Code and models are publicly available.

Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan
arXiv:2501.17811 · cs.AI, cs.CL, cs.CV · submitted Jan 29, 2025
abstract · pdf · html · Research paper. arXiv admin note: text overlap with arXiv:2410.13848

add comment on HN