about
Cellular automata as convolutional neural networks (arxiv.org)
83 points by wil3 on Sep 30, 2018 | hide | past | pdf | 6 comments on HN

In plain words: A cellular automaton is a grid where each cell updates by a fixed rule from its neighbors; any such rule can be written as a convolutional network that learns it from video. Simple rule tables gave networks layered, specialized structure; complex rules gave shallower ones.

Abstract

Deep learning techniques have recently demonstrated broad success in predicting complex dynamical systems ranging from turbulence to human speech, motivating broader questions about how neural networks encode and represent dynamical rules. We explore this problem in the context of cellular automata (CA), simple dynamical systems that are intrinsically discrete and thus difficult to analyze using standard tools from dynamical systems theory. We show that any CA may readily be represented using a convolutional neural network with a network-in-network architecture. This motivates our development of a general convolutional multilayer perceptron architecture, which we find can learn the dynamical rules for arbitrary CA when given videos of the CA as training data. In the limit of large network widths, we find that training dynamics are nearly identical across replicates, and that common patterns emerge in the structure of networks trained on different CA rulesets. We train ensembles of networks on randomly-sampled CA, and we probe how the trained networks internally represent the CA rules using an information-theoretic technique based on distributions of layer activation patterns. We find that CA with simpler rule tables produce trained networks with hierarchical structure and layer specialization, while more complex CA produce shallower representations---illustrating how the underlying complexity of the CA's rules influences the specificity of these internal representations. Our results suggest how the entropy of a physical process can affect its representation when learned by neural networks.

William Gilpin
arXiv:1809.02942 · nlin.CG, cond-mat.dis-nn, cs.NE, physics.comp-ph · submitted Sep 9, 2018 · updated Jan 16, 2020
abstract · pdf · html · 8 pages, 4 figures (+Appendix)

add comment on HN
Also discussed: Aug 2020 (87 points, 14 comments)

Nice try, Stephen Wolfram...
Isn't Wolfram more likely to do convolutional neural networks as cellular automata?
Uh oh, I've been exposed
Anything surprising here, if anyone’s up to date on this stuff? I thought it’s a given that since both CNNs and CAs are Turing complete, you can use one to simulate the other.
> I thought it’s a given that since both CNNs and CAs are Turing complete, you can use one to simulate the other.

They are both Turing complete, yes, but the similarities are deeper. Recurrent CNNs are a very natural way to describe CAs because they both impose exactly the same prior: that the system's underlying laws are spatially quantized, local, and translationally invariant.

There hasn't been much work done in this area. People have made the connection before (e.g. this paper from Sutskever: https://arxiv.org/pdf/1511.08228.pdf) but AFAIK this is the first work to actually look deeply into how well you can recover CA rules with neural nets and gradient descent.

If we are to believe Steven Wolfram's proposition that the universe is a cellular automaton at its core, then learning the rule by gradient descent seems pretty obviously useful, and not something that anyone has really tried yet.

To be more explicit: applying gradient descent to cellular automata has a reasonable likelihood to being the only way to uncover the fundamental laws of reality, and essentially nobody is working on it except OP. So it's a useful paper!

Turing completeness is often too coarse-grained of a view to be meaningful. The interesting thing here is that convolutional nets and cellular automata perform a similar kind of spatially partitioned computation, so connections between the two may be enlightening.