In plain words: A neural network learns to command small StarCraft armies from raw game numbers, by trying out whole battle plans and training on the best ones. It found real tactics for armies of up to 15 units, where standard trial-and-error learners failed.
Abstract · Episodic Exploration for Deep Deterministic Policies: An Application to StarCraft Micromanagement Tasks
We consider scenarios from the real-time strategy game StarCraft as new benchmarks for reinforcement learning algorithms. We propose micromanagement tasks, which present the problem of the short-term, low-level control of army members during a battle. From a reinforcement learning point of view, these scenarios are challenging because the state-action space is very large, and because there is no obvious feature representation for the state-action evaluation function. We describe our approach to tackle the micromanagement scenarios with deep neural network controllers from raw state features given by the game engine. In addition, we present a heuristic reinforcement learning algorithm which combines direct exploration in the policy space and backpropagation. This algorithm allows for the collection of traces for learning using deterministic policies, which appears much more efficient than, for example, ε-greedy exploration. Experiments show that with this algorithm, we successfully learn non-trivial strategies for scenarios with armies of up to 15 agents, where both Q-learning and REINFORCE struggle.
Nicolas Usunier, Gabriel Synnaeve, Zeming Lin, Soumith Chintala
arXiv:1609.02993 · cs.AI, cs.LG · submitted Sep 10, 2016 · updated Nov 26, 2016
abstract · pdf · html · 18 pages, 1 figure (2 plots), 2 tables
So perhaps it's worth pointing out, that this paper specifically addresses a sub-problem of Starcraft play, micromanagement ('micro') [1]
The game engine runs at 24 frames per second. (As an aside, 'frames' in this context likely does not map to physical FPS of the display).
>We ran all the following experiments with a skip_frames of 9 (meaning that we take about 2.6 actions per unit per second).
The research team found that attempting to move at a superhuman pace (eg one action every frame), resulted in a subpar performance and hyper-parameterization indicated 2.6 to be an ideal action per second.
In context, this translates to an APM of 156. Or, roughly half that of professional Korean e-athletes. [2]
[1] https://en.wikipedia.org/wiki/Micromanagement_(gameplay)
[2] https://en.wikipedia.org/wiki/Actions_per_minute