In plain words: A language model is trained on curated text-instruction-to-layout examples so it can turn a plain description into a ready-made poster or interface design. It beat the strongest competing systems, including ones built on the best commercial chatbots, on image generation and design tests.
Abstract
Automatic generation of graphical layouts is crucial for many real-world applications, including designing posters, flyers, advertisements, and graphical user interfaces. Given the incredible ability of Large language models (LLMs) in both natural language understanding and generation, we believe that we could customize an LLM to help people create compelling graphical layouts starting with only text instructions from the user. We call our method TextLap (text-based layout planning). It uses a curated instruction-based layout planning dataset (InsLap) to customize LLMs as a graphic designer. We demonstrate the effectiveness of TextLap and show that it outperforms strong baselines, including GPT-4 based methods, for image generation and graphical design benchmarks.
Jian Chen, Ruiyi Zhang, Yufan Zhou, Jennifer Healey, Jiuxiang Gu, Zhiqiang Xu, Changyou Chen
arXiv:2410.12844 · cs.CL, cs.LG · submitted Oct 9, 2024
abstract · pdf · html · Accepted to the EMNLP Findings