about
Mental time travel: Hierarchical memory for reinforcement learning agents (2021) (arxiv.org)
82 points by kordlessagain on May 26, 2022 | hide | past | pdf | 8 comments on HN

In plain words: The agent stores its past in chunks, scans summaries to pick the right one, then reads it closely to recall details without replaying every step. It beat other memory designs at remembering where an object was hidden, even after much longer delays than in training.

Abstract · Towards mental time travel: a hierarchical memory for reinforcement learning agents

Reinforcement learning agents often forget details of the past, especially after delays or distractor tasks. Agents with common memory architectures struggle to recall and integrate across multiple timesteps of a past event, or even to recall the details of a single timestep that is followed by distractor tasks. To address these limitations, we propose a Hierarchical Chunk Attention Memory (HCAM), which helps agents to remember the past in detail. HCAM stores memories by dividing the past into chunks, and recalls by first performing high-level attention over coarse summaries of the chunks, and then performing detailed attention within only the most relevant chunks. An agent with HCAM can therefore "mentally time-travel" -- remember past events in detail without attending to all intervening events. We show that agents with HCAM substantially outperform agents with other memory architectures at tasks requiring long-term recall, retention, or reasoning over memory. These include recalling where an object is hidden in a 3D environment, rapidly learning to navigate efficiently in a new neighborhood, and rapidly learning and retaining new object names. Agents with HCAM can extrapolate to task sequences much longer than they were trained on, and can even generalize zero-shot from a meta-learning setting to maintaining knowledge across episodes. HCAM improves agent sample efficiency, generalization, and generality (by solving tasks that previously required specialized architectures). Our work is a step towards agents that can learn, interact, and adapt in complex and temporally-extended environments.

Andrew Kyle Lampinen, Stephanie C. Y. Chan, Andrea Banino, Felix Hill
arXiv:2105.14039 · cs.LG, cs.AI, cs.NE · submitted May 28, 2021 · updated Dec 8, 2021
abstract · pdf · html · NeurIPS 2021; 10 pages main text; 29 pages total

add comment on HN

Fun coincidence, I read this paper as part of my AI research just recently. It's actually quite neat! Mostly because there is really not much like it in the sense that it focuses on exploring a memory mechanism for RL agents - most of the time people just stack frames or slap on an RNN and call it a day. The set of tasks evaluated on is cool too.
> . HCAM stores memories by dividing the past into chunks, and recalls by first performing high-level attention over coarse summaries of the chunks, and then performing detailed attention within only the most relevant chunks.

This sounds just like the techniques the human 'memory athletics' gurus teach.

The original gurus or the fake ones after everyone realized it was profitable?
I think I first heard about chunking from one of the real ones trying to spin a profit.
> or reasoning over memory

Where is the "reasoning" in this? sounds like pattern recognition to me.

The authors are seemingly falling into the fallacy of replicating the human experience, which simply leads to ultimately inefficient Alt Intelligence. While an interesting idea which could have a couple very rare useful use-cases, this seems like another of the million [Deep Learning + a minor wacky idea] research papers.
I have not read the paper, but for some high level context, it's written by a team of researchers at DeepMind, and was presented at NeurIPS, so I'm going to go out on a limb and give the authors the benefit of the doubt that they probabally have some idea of what they're talking about.
Can you explain why it's a mistake in this particular case?