In plain words: A text-to-image system that starts with random noise and gradually shapes it into a picture matching the prompt. In side-by-side comparisons, people picked its images more often than those from the other leading generators of the time.
Abstract · Imagen 3
We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred over other state-of-the-art (SOTA) models at the time of evaluation. In addition, we discuss issues around safety and representation, as well as methods we used to minimize the potential harm of our models.
Imagen-Team-Google, :, Jason Baldridge, Jakob Bauer, Mukul Bhutani, Nicole Brichtova, Andrew Bunner, Lluis Castrejon, Kelvin Chan, Yichang Chen, Sander Dieleman, Yuqing Du, et al.
arXiv:2408.07009 · cs.CV · submitted Aug 13, 2024 · updated Dec 21, 2024
abstract · pdf · html