about
Benchmark to measure AI on graphic design tasks (arxiv.org)
5 points by purvanshi 179 days ago | hide | past | pdf | 2 comments on HN

In plain words: A new test set of 50 professional graphic-design tasks, from layouts and typography to infographics and animation, checks whether AI can understand and create design work. Top models grasp the idea but stumble on precise placement, valid vector drawings, fine typography, and animation timing.

Abstract · Graphic-Design-Bench: A Comprehensive Benchmark for Evaluating AI on Graphic Design Tasks

We introduce GraphicDesignBench (GDB), the first comprehensive benchmark suite designed specifically to evaluate AI models on the full breadth of professional graphic design tasks. Unlike existing benchmarks that focus on natural-image understanding or generic text-to-image synthesis, GDB targets the unique challenges of professional design work: translating communicative intent into structured layouts, rendering typographically faithful text, manipulating layered compositions, producing valid vector graphics, and reasoning about animation. The suite comprises 50 tasks organized along five axes: layout, typography, infographics, template & design semantics and animation, each evaluated under both understanding and generation settings, and grounded in real-world design templates drawn from the LICA layered-composition dataset. We evaluate a set of frontier closed-source models using a standardized metric taxonomy covering spatial accuracy, perceptual quality, text fidelity, semantic alignment, and structural validity. Our results reveal that current models fall short on the core challenges of professional design: spatial reasoning over complex layouts, faithful vector code generation, fine-grained typographic perception, and temporal decomposition of animations remain largely unsolved. While high-level semantic understanding is within reach, the gap widens sharply as tasks demand precision, structure, and compositional awareness. GDB provides a rigorous, reproducible testbed for tracking progress toward AI systems that can function as capable design collaborators. The full evaluation framework is publicly available.

Adrienne Deganutti, Elad Hirsch, Haonan Zhu, Jaejung Seol, Purvanshi Mehta
arXiv:2604.04192 · cs.CV, cs.AI, cs.LG · submitted Apr 5, 2026 · updated Apr 7, 2026
abstract · pdf · html

add comment on HN
Also discussed: Apr 2026 (8 points, 0 comments)

Single layout generation is hard. Generating template variants, i.e., multiple layouts that share a structure but differ in style, color, and content, is a completely different problem.

We tested style completion and recoloring on template families. Structural fidelity is high (position and area preservation near 100% for structural generation), but palette coverage lags at 77.6%. Worse, SSIM and LPIPS actively mislead: a structurally valid, style-consistent output scores lower than a hallucinated one that happens to agree more on pixels.

The take away is that pixel metrics are the wrong evaluation substrate for design. The field needs structure-aware metrics that operate on extracted primitives such as bounding boxes, color tokens, font properties instead of raw pixels.

We asked frontier models to detect components in a graphic design layout. The best result: 6.4% [email protected]. For context, natural image detection benchmarks sit above 60%. We're an order of magnitude behind on a task every professional design tool does trivially. So while models can talk about design, they still struggle to locate it. This means the system lacks the spatial grounding needed for reliable editing, completion, or structure-aware manipulation.