about
Toward an AI Physicist for Unsupervised Learning (arxiv.org)
111 points by rcshubhadeep on Nov 5, 2018 | hide | past | pdf | 5 comments on HN

In plain words: Instead of a network learning everything, this agent learns theories—formulas that predict motion and know where they apply—and merges them in a hub. On worlds with gravity and bounces, its errors were about a billion times smaller than a standard neural net's.

Abstract

We investigate opportunities and challenges for improving unsupervised machine learning using four common strategies with a long history in physics: divide-and-conquer, Occam's razor, unification and lifelong learning. Instead of using one model to learn everything, we propose a novel paradigm centered around the learning and manipulation of *theories*, which parsimoniously predict both aspects of the future (from past observations) and the domain in which these predictions are accurate. Specifically, we propose a novel generalized-mean-loss to encourage each theory to specialize in its comparatively advantageous domain, and a differentiable description length objective to downweight bad data and "snap" learned theories into simple symbolic formulas. Theories are stored in a "theory hub", which continuously unifies learned theories and can propose theories when encountering new environments. We test our implementation, the toy "AI Physicist" learning agent, on a suite of increasingly complex physics environments. From unsupervised observation of trajectories through worlds involving random combinations of gravity, electromagnetism, harmonic motion and elastic bounces, our agent typically learns faster and produces mean-squared prediction errors about a billion times smaller than a standard feedforward neural net of comparable complexity, typically recovering integer and rational theory parameters exactly. Our agent successfully identifies domains with different laws of motion also for a nonlinear chaotic double pendulum in a piecewise constant force field.

Tailin Wu, Max Tegmark
arXiv:1810.10525 · physics.comp-ph, cond-mat.dis-nn, cs.LG · submitted Oct 24, 2018 · updated Sep 2, 2019
abstract · pdf · html · Replaced to match accepted PRE version. Added references, improved discussion. 22 pages, 7 figs

add comment on HN
Also discussed: Nov 2020 (10 points, 0 comments)

The computer program BACON (1987) of Nobel Prize winner Herbert Simon was given the distances of planets from the sun together with their period of revolution and it independently rediscovered Kepler's third law, illustrating how far the positivism at work in "AI" can go. But Kepler's achievement was not determining a - straightforward - relation between two rows of numbers: it was to figure out which numbers should be related, and Kepler's real achievement was actually finding the right question. Incidentally, Kepler stated a fourth law relating planets to perfect polyhedra, and one wonders why this fourth law has not been rediscovered by computer yet...

-- Jean-Yves Girard: Locus Solum

Bacon was the dissertation project of Pat Langley under Herbert Simon. It was one of several dissertations exploring rational reconstruction of previous discoveries. https://www.ijcai.org/Proceedings/81-1/Papers/025.pdf
Perhaps BACON would have discovered Kepler's 4th law if the AI was programmed by mystics? :)
So I'm still trying to grok this, but this could be a very important result if it could be generalized to the problem of meta-learning and model selection.

In that setting, the model selection oracle could access a shared knowledge pool of learned theorems that are biases about what kinds of models are better. e.g. convolutional nets are better than fully connected nets for vision tasks.

This could break us out of the diminishing returns we have seen with deep learning, by allowing us to better explore the space of compact model architectures, and develop shared biases about what is better. For example, learning the programmatic generation of Inception-like networks.

Bonus points if you want to add a blockchain connection, to decentralize the accumulation of the shared theorem base: Proof of work is figuring out what biases are true over some benchmark, which can be stored to a distributed ledger. Competing annotations lead malicious and noisy results to be penalized.

I got introduced to Max Tegmark's work last year. Got his book - Life 3.0 and Our Mathematical Universe