In plain words: A model for sorting images is built from a chain of small number blocks that together compress a huge table of possible patterns, a trick borrowed from physics. On handwritten digits it misclassified under 1% of test images.
Abstract
Tensor networks are efficient representations of high-dimensional tensors which have been very successful for physics and mathematics applications. We demonstrate how algorithms for optimizing such networks can be adapted to supervised learning tasks by using matrix product states (tensor trains) to parameterize models for classifying images. For the MNIST data set we obtain less than 1% test set classification error. We discuss how the tensor network form imparts additional structure to the learned model and suggest a possible generative interpretation.
E. Miles Stoudenmire, David J. Schwab
arXiv:1605.05775 · stat.ML, cond-mat.str-el, cs.LG · submitted May 18, 2016 · updated May 18, 2017
abstract · pdf · html · 11 pages, 15 figures; updated version includes corrections, links to sample codes, expanded discussion, and additional references