about
A Deep Hierarchical Approach to Lifelong Learning in Minecraft (arxiv.org)
46 points by jonbaer on Dec 30, 2016 | hide | past | pdf | 7 comments on HN

In plain words: The system learns reusable skills in Minecraft and squeezes them into one network, keeping old skills while picking up new tasks instead of relearning from scratch. In Minecraft sub-tasks, it beat the usual network that learns each task alone, using fewer training examples.

Abstract

We propose a lifelong learning system that has the ability to reuse and transfer knowledge from one task to another while efficiently retaining the previously learned knowledge-base. Knowledge is transferred by learning reusable skills to solve tasks in Minecraft, a popular video game which is an unsolved and high-dimensional lifelong learning problem. These reusable skills, which we refer to as Deep Skill Networks, are then incorporated into our novel Hierarchical Deep Reinforcement Learning Network (H-DRLN) architecture using two techniques: (1) a deep skill array and (2) skill distillation, our novel variation of policy distillation (Rusu et. al. 2015) for learning skills. Skill distillation enables the HDRLN to efficiently retain knowledge and therefore scale in lifelong learning, by accumulating knowledge and encapsulating multiple reusable skills into a single distilled network. The H-DRLN exhibits superior performance and lower learning sample complexity compared to the regular Deep Q Network (Mnih et. al. 2015) in sub-domains of Minecraft.

Chen Tessler, Shahar Givony, Tom Zahavy, Daniel J. Mankowitz, Shie Mannor
arXiv:1604.07255 · cs.AI, cs.LG · submitted Apr 25, 2016 · updated Nov 30, 2016
abstract · pdf · html

add comment on HN

And sadly, another paper about software with no source code whatsoever.
Email them politely asking for the code.
Doesn't scale usefully nor temporally.
It scales better than complaining on Hacker News.
It does effectively warn the readership and saves them from having to trawl the document and the web on their own for the source.
This looks like a really cool architecture, but I'm mildly uncertain about how well it scales given the simplicity of the tasks demonstrated. Delivery tasks are the most basic tasks in RL research. I'd be interested in predator/hazard avoidance, as well as some sort of rudimentary reasoning (might require analogy structure built-in?).

Additionally, the pure-pixel input is neato, but the discrete outputs seem limiting.

That being said, I'm overall very happy with the direction of research!