about
E-HBA: Using Action Policies for Expert Advice and Agent Typification (arxiv.org)
2 points by sel1 on Jul 25, 2019 | hide | past | pdf | discuss on HN

In plain words: It uses a fixed set of play styles two ways: as experts to copy when picking your moves, and as types to guess the opponent's move. Blending each expert's past score with a predicted future score, it improved standard expert-following strategies in repeated games.

Abstract

Past research has studied two approaches to utilise predefined policy sets in repeated interactions: as experts, to dictate our own actions, and as types, to characterise the behaviour of other agents. In this work, we bring these complementary views together in the form of a novel meta-algorithm, called Expert-HBA (E-HBA), which can be applied to any expert algorithm that considers the average (or total) payoff an expert has yielded in the past. E-HBA gradually mixes the past payoff with a predicted future payoff, which is computed using the type-based characterisation. We present results from a comprehensive set of repeated matrix games, comparing the performance of several well-known expert algorithms with and without the aid of E-HBA. Our results show that E-HBA has the potential to significantly improve the performance of expert algorithms.

Stefano V. Albrecht, Jacob W. Crandall, Subramanian Ramamoorthy
arXiv:1907.09810 · cs.AI, cs.MA · submitted Jul 23, 2019
abstract · pdf · html · Proceedings of the Second Workshop on Multiagent Interaction without Prior Coordination (MIPC), 2015

add comment on HN