about
Learning the Arrow of Time (arxiv.org)
2 points by sel1 on Jul 3, 2019 | hide | past | pdf | discuss on HN

In plain words: The system learns which situations tend to come later rather than earlier in an environment, giving it a built-in sense of time's direction. That signal can measure reachability and flag side effects, and it matched a known physics-based measure of time's arrow.

Abstract

We humans seem to have an innate understanding of the asymmetric progression of time, which we use to efficiently and safely perceive and manipulate our environment. Drawing inspiration from that, we address the problem of learning an arrow of time in a Markov (Decision) Process. We illustrate how a learned arrow of time can capture meaningful information about the environment, which in turn can be used to measure reachability, detect side-effects and to obtain an intrinsic reward signal. We show empirical results on a selection of discrete and continuous environments, and demonstrate for a class of stochastic processes that the learned arrow of time agrees reasonably well with a known notion of an arrow of time given by the celebrated Jordan-Kinderlehrer-Otto result.

Nasim Rahaman, Steffen Wolf, Anirudh Goyal, Roman Remme, Yoshua Bengio
arXiv:1907.01285 · cs.LG, cs.AI · submitted Jul 2, 2019
abstract · pdf · html · A shorter version of this work was presented at the Theoretical Phyiscs for Deep Learning Workshop, ICML 2019

add comment on HN