In plain words: Instead of predicting frames, this model guesses each pixel's color in turn, following the video's order in time, space, and color. On Moving MNIST it nearly hit the best possible score, far above the previous best, with videos barely different from real ones.
Abstract
We propose a probabilistic video model, the Video Pixel Network (VPN), that estimates the discrete joint distribution of the raw pixel values in a video. The model and the neural architecture reflect the time, space and color structure of video tensors and encode it as a four-dimensional dependency chain. The VPN approaches the best possible performance on the Moving MNIST benchmark, a leap over the previous state of the art, and the generated videos show only minor deviations from the ground truth. The VPN also produces detailed samples on the action-conditional Robotic Pushing benchmark and generalizes to the motion of novel objects.
Nal Kalchbrenner, Aaron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, Koray Kavukcuoglu
arXiv:1610.00527 · cs.CV, cs.LG · submitted Oct 3, 2016
abstract · pdf · html · 16 pages