about
Incentivising cooperation by rewarding the weakest member (arxiv.org)
2 points by PaulHoule on Dec 4, 2022 | hide | past | pdf | discuss on HN

In plain words: Groups of computer agents are scored by how their weakest member performs, rather than each agent chasing its own reward. Compared with rewarding each agent for its own results, this made behavior fairer and more even while still improving every agent's outcome.

Abstract

Autonomous agents that act with each other on behalf of humans are becoming more common in many social domains, such as customer service, transportation, and health care. In such social situations greedy strategies can reduce the positive outcome for all agents, such as leading to stop-and-go traffic on highways, or causing a denial of service on a communications channel. Instead, we desire autonomous decision-making for efficient performance while also considering equitability of the group to avoid these pitfalls. Unfortunately, in complex situations it is far easier to design machine learning objectives for selfish strategies than for equitable behaviors. Here we present a simple way to reward groups of agents in both evolution and reinforcement learning domains by the performance of their weakest member. We show how this yields ``fairer'' more equitable behavior, while also maximizing individual outcomes, and we show the relationship to biological selection mechanisms of group-level selection and inclusive fitness theory.

Jory Schossau, Bamshad Shirmohammadi, Arend Hintze
arXiv:2212.00119 · cs.MA, cs.AI, cs.LG · submitted Oct 4, 2022
abstract · pdf · html · 11 pages, 4 figures

add comment on HN