about
Logo Synthesis and Manipulation with Clustered Generative Adversarial Networks (arxiv.org)
10 points by relate on Dec 14, 2017 | hide | past | pdf | 8 comments on HN

In plain words: They gathered over 600,000 web logos and trained an image generator that competes against a checker, adding labels from grouping similar images so it doesn't just churn out the same design. It produced varied, believable logos and set top image-quality scores on a standard test set.

Abstract

Designing a logo for a new brand is a lengthy and tedious back-and-forth process between a designer and a client. In this paper we explore to what extent machine learning can solve the creative task of the designer. For this, we build a dataset -- LLD -- of 600k+ logos crawled from the world wide web. Training Generative Adversarial Networks (GANs) for logo synthesis on such multi-modal data is not straightforward and results in mode collapse for some state-of-the-art methods. We propose the use of synthetic labels obtained through clustering to disentangle and stabilize GAN training. We are able to generate a high diversity of plausible logos and we demonstrate latent space exploration techniques to ease the logo design task in an interactive manner. Moreover, we validate the proposed clustered GAN training on CIFAR 10, achieving state-of-the-art Inception scores when using synthetic labels obtained via clustering the features of an ImageNet classifier. GANs can cope with multi-modal data by means of synthetic labels achieved through clustering, and our results show the creative potential of such techniques for logo synthesis and manipulation. Our dataset and models will be made publicly available at https://data.vision.ee.ethz.ch/cvl/lld/.

Alexander Sage, Eirikur Agustsson, Radu Timofte, Luc Van Gool
arXiv:1712.04407 · cs.CV, cs.LG, stat.ML · submitted Dec 12, 2017
abstract · pdf · html

add comment on HN

This looks interesting, but the results are somewhat blurry.

Has anybody ever tried to use features of the logos (number of shapes, shape size, position, color, curvature, shape parents/children, etc.) instead of raw pixel data to train GANs?

Here comes the shameless plug :

ImageTracer is a simple raster image tracer and vectorizer that outputs SVG, 100% free, Public Domain.

Available in JavaScript (works both in the browser and with Node.js), "desktop" Java and "Android" Java:

https://github.com/jankovicsandras/imagetracerjs

https://github.com/jankovicsandras/imagetracerjava

https://github.com/jankovicsandras/imagetracerandroid

I think logos are a tough problem for convnets because they're not very compositional - ie. they're not made of heirarchically nested parts.

the space of logos is also probably not continuous - eg. there is a logo in the latent space between nike and apple, but it's unlikely to be aesthetic.

Nipple? Apik? Could be a sportswatch. There might be a few clusters along the path through latent space (if that makes any sense, just skimmed the paper mostly for the figures).

The attempt here seems to be really naive, I agree. But why are logos not compositing? Coat of arms are frequently described in such a manner that would allow to mix them. But then, the traditional artistic combinations of different ones into new are not mere half way morphs. And a classic logo needs to be compositional, because it's easier to perceive (decompose), e.g. hammer and sickle. Scientific Icons are frequently using mathematical patterns and plots, which tickle the eye in quite a different manner. I thought the nike swoosh comes from that rather abstract direction, whereas the apple is quite objective. Both are pictographs, but only the apple is a logo (from logos, ie. speaking).

there are logos that are compositional, but most logos are abstractive - eg. the hammer and sickle logo only makes sense because we have prior knowledge of what hammers and sickles look like. You could learn an abstract representation of hammers from a dataset of hammers, but not from a dataset of logos.

I think GANs work best on images with hierarchical composition like human faces.

These are iconographs, in the strict sense, not the full logos. The different google Gs don't really speak for themselves. The Y combinator Y is really not distinctive, either. The first few figures show fav-icons, I'd thought.
Hi, I'm one of the authors. That is correct in a strict sense, but we wanted to focus on the more 'creative' part of logos rather than the text. GANs are known to struggle with high resolutions, but we note that we show higher resolution logos later in the paper (see page 12 and 15 e.g) which is trained on the smaller but higher-resolution version of our dataset.
This title should be changed to "Smudge Synthesis". Move along nothing to see here. Actually the dataset of 600k logos is probably interesting. I bet someone who had some time could do a hugely better job.
I found interesting the identification of different clusters. Now do some learning over the clusters. Letter cut out from colored shape seems to be the most prominent feature, and the most boring one. But some of the shapes (without cut-outs) are rather interesting. Some of the figures from he article could become icons themselves, reminiscent of scientific plots (cf. https://upload.wikimedia.org/wikipedia/commons/thumb/b/b0/Un...).

The problem is, a logo should be as unique as possible, so mechanical derivatives aren't convincing