about
Five-Dollar Model: Generating Game Maps and Sprites from Sentence Embeddings (arxiv.org)
1 point by PaulHoule on Aug 14, 2023 | hide | past | pdf | discuss on HN

In plain words: A tiny text-to-image generator turns a sentence into pixel-art game maps, sprites, or emoji, using new tricks to stretch limited training data. Despite its small size and few examples, its images keep the prompt's meaning and look good, unlike big systems needing huge datasets.

Abstract · The Five-Dollar Model: Generating Game Maps and Sprites from Sentence Embeddings

The five-dollar model is a lightweight text-to-image generative architecture that generates low dimensional images from an encoded text prompt. This model can successfully generate accurate and aesthetically pleasing content in low dimensional domains, with limited amounts of training data. Despite the small size of both the model and datasets, the generated images are still able to maintain the encoded semantic meaning of the textual prompt. We apply this model to three small datasets: pixel art video game maps, video game sprite images, and down-scaled emoji images and apply novel augmentation strategies to improve the performance of our model on these limited datasets. We evaluate our models performance using cosine similarity score between text-image pairs generated by the CLIP VIT-B/32 model.

Timothy Merino, Roman Negri, Dipika Rajesh, M Charity, Julian Togelius
arXiv:2308.04052 · cs.LG, cs.CL, cs.CV · submitted Aug 8, 2023
abstract · pdf · html · to be published in AIIDE 2023

add comment on HN
Also discussed: Aug 2023 (32 points, 2 comments)