In plain words: AIXI, a theoretical agent that can learn any computable world, has probabilities that don't add to one; that gap is read here as its chance of dying. It proves the reading works, and finds rescaling rewards flips the agent from suicidal to stubbornly self-preserving.
Abstract
Reinforcement learning (RL) is a general paradigm for studying intelligent behaviour, with applications ranging from artificial intelligence to psychology and economics. AIXI is a universal solution to the RL problem; it can learn any computable environment. A technical subtlety of AIXI is that it is defined using a mixture over semimeasures that need not sum to 1, rather than over proper probability measures. In this work we argue that the shortfall of a semimeasure can naturally be interpreted as the agent's estimate of the probability of its death. We formally define death for generally intelligent agents like AIXI, and prove a number of related theorems about their behaviour. Notable discoveries include that agent behaviour can change radically under positive linear transformations of the reward signal (from suicidal to dogmatically self-preserving), and that the agent's posterior belief that it will survive increases over time.
Jarryd Martin, Tom Everitt, Marcus Hutter
arXiv:1606.00652 · cs.AI · submitted Jun 2, 2016
abstract · pdf · html · Conference: Artificial General Intelligence (AGI) 2016 13 pages, 2 figures
Basically this is saying the reinforcement learning systems (and other ML systems?) can get stuck in states from which they cannot escape. They label these states "death"
They go on to define various agent behaviours, which given starting states lead to different equilibriums. The ones that actively lead to certain death they label "suicide"