In plain words: A system colors black-and-white photos automatically by guessing a color for each part, trained on over a million color images and nudged toward brighter, more varied hues. People mistook its colorized photos for real ones 32% of the time, more often than with earlier tools.
Abstract
Given a grayscale photograph as input, this paper attacks the problem of hallucinating a plausible color version of the photograph. This problem is clearly underconstrained, so previous approaches have either relied on significant user interaction or resulted in desaturated colorizations. We propose a fully automatic approach that produces vibrant and realistic colorizations. We embrace the underlying uncertainty of the problem by posing it as a classification task and use class-rebalancing at training time to increase the diversity of colors in the result. The system is implemented as a feed-forward pass in a CNN at test time and is trained on over a million color images. We evaluate our algorithm using a "colorization Turing test," asking human participants to choose between a generated and ground truth color image. Our method successfully fools humans on 32% of the trials, significantly higher than previous methods. Moreover, we show that colorization can be a powerful pretext task for self-supervised feature learning, acting as a cross-channel encoder. This approach results in state-of-the-art performance on several feature learning benchmarks.
Richard Zhang, Phillip Isola, Alexei A. Efros
arXiv:1603.08511 · cs.CV · submitted Mar 28, 2016 · updated Oct 5, 2016
abstract · pdf · html
And later in the article: "However, the results from these and other past attempts tend to look desaturated. One explanation is that [1,2] use loss functions that encourage conservative predictions. [...] We instead utilize a loss tailored to the colorization problem. As pointed out by [3], color prediction is inherently multimodal – many objects, such as a shirt, can plausibly be colored one of several distinct values."
This suggests that adversarial training might be a good fit. It transforms the goal from "reproduce the original colors", which is in general impossible, to "produce as convincing a colorization as possible", which is the real goal.