about
Qwen-Image-Flash: Beyond Objective Design (arxiv.org)
2 points by gmays 116 days ago | hide | past | pdf | discuss on HN

In plain words: They reworked how a big text-to-image and image-editing model is squeezed into one that draws in a few steps, testing different data mixes, teacher hints, and task blends. The results show these recipe choices matter as much as the training goal, which past work ignored.

Abstract · Qwen-Image-Flash: Rethinking the Training Recipe for Few-Step Distillation

Few-step distillation has emerged as a critical component in the development of advanced visual generative foundation models, substantially reducing inference overhead while enabling real-time generation and cost-efficient deployment across a broad range of practical scenarios. However, prior work has predominantly focused on advancing training objectives, while comparatively overlooking the training recipe, which has become increasingly critical in the era of large-scale foundation models. In this work, we systematically revisit the training recipe under the well-established distribution matching distillation (DMD) framework for both text-to-image generation and image editing, focusing on three key dimensions: training data composition, teacher guidance within DMD, and task mixture. Our empirical analysis reveals several non-obvious and counterintuitive phenomena, ultimately motivating the development of Qwen-Image-Flash. These findings highlight that effective few-step distillation depends not only on carefully designed objectives, but also on a principled training recipe.

Tianhe Wu, Zikai Zhou, Kun Yan, Kaiyuan Gao, Lihan Jiang, Jiahao Li, Jie Zhang, Ningyuan Tang, Shengming Yin, Xiaoyue Chen, Xiao Xu, Yilei Chen, et al.
arXiv:2606.03746 · cs.CV, cs.AI, cs.GR, cs.LG · submitted Jun 2, 2026 · updated Sep 14, 2026
abstract · pdf · html

add comment on HN