about
Ranking of popular image generation AI models (incl. Flux) from 2M votes (arxiv.org)
1 point by maalber on Sep 20, 2024 | hide | past | pdf | 1 comment on HN

In plain words: A system that gathers preference votes from a large, worldwide crowd to judge AI-made images on looks, sense, and prompt match. It collected over 2 million votes to rank four image generators, and the crowd's mix of people mirrors the world, unlike judging panels.

Abstract · Finding the Subjective Truth: Collecting 2 Million Votes for Comprehensive Gen-AI Model Evaluation

Efficiently evaluating the performance of text-to-image models is difficult as it inherently requires subjective judgment and human preference, making it hard to compare different models and quantify the state of the art. Leveraging Rapidata's technology, we present an efficient annotation framework that sources human feedback from a diverse, global pool of annotators. Our study collected over 2 million annotations across 4,512 images, evaluating four prominent models (DALL-E 3, Flux.1, MidJourney, and Stable Diffusion) on style preference, coherence, and text-to-image alignment. We demonstrate that our approach makes it feasible to comprehensively rank image generation models based on a vast pool of annotators and show that the diverse annotator demographics reflect the world population, significantly decreasing the risk of biases.

Dimitrios Christodoulou, Mads Kuhlmann-Jørgensen
arXiv:2409.11904 · cs.CV, cs.AI · submitted Sep 18, 2024 · updated Oct 15, 2024
abstract · pdf · html

add comment on HN

Efficiently evaluating the performance of text-to-image models is difficult as it inherently requires subjective judgment and human preference, making it hard to compare different models and ultimately quantify progress. We developed a system to efficiently source annotations at scale and based on this present a framework for rigorous evaluation of image generation models. With this, we present the ranking of four popular image generation models, MidJourney, DALLE-3, Stable Diffusions, and the latest star, Flux.1. The ranking is based on more than 2 million annotations across 4512 images and three criterias; style, coherence, and text-to-image alignment. Through integration with mobile apps, we reach a diverse set of annotators and we show that the regional distribution closely resembles the distribution of the world population ensuring lower risk of biases.

If you want to get a quick feel of our system, check out our free Compare Tool (https://www.rapidata.ai/compare) which resembles a single one of the 27k comparisons created for the ranking.