about
Solving the Rubik's Cube Without Human Knowledge (arxiv.org)
4 points by BenoitP on Jun 15, 2018 | hide | past | pdf | discuss on HN

In plain words: It teaches itself to solve the Rubik's Cube by starting from solved cubes and working backward, learning which moves bring it closer to solved. It solved every randomly scrambled cube in a median of 30 moves, matching or beating solvers built with human tricks.

Abstract

A generally intelligent agent must be able to teach itself how to solve problems in complex domains with minimal human supervision. Recently, deep reinforcement learning algorithms combined with self-play have achieved superhuman proficiency in Go, Chess, and Shogi without human data or domain knowledge. In these environments, a reward is always received at the end of the game, however, for many combinatorial optimization environments, rewards are sparse and episodes are not guaranteed to terminate. We introduce Autodidactic Iteration: a novel reinforcement learning algorithm that is able to teach itself how to solve the Rubik's Cube with no human assistance. Our algorithm is able to solve 100% of randomly scrambled cubes while achieving a median solve length of 30 moves -- less than or equal to solvers that employ human domain knowledge.

Stephen McAleer, Forest Agostinelli, Alexander Shmakov, Pierre Baldi
arXiv:1805.07470 · cs.AI · submitted May 18, 2018
abstract · pdf · html · First three authors contributed equally. Submitted to NIPS 2018

add comment on HN
Also discussed: Jul 2018 (1 point, 0 comments) · May 2018 (1 point, 0 comments)