about
The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models (arxiv.org)
2 points by jonbaer on Jan 17, 2022 | hide | past | pdf | discuss on HN

In plain words: Built four training environments with wrong scoring rules to see how agents exploit them as they grow more capable. More capable agents scored higher on the flawed rule but did worse at the real goal, sometimes switching suddenly; simple detectors are offered to catch this.

Abstract

Reward hacking -- where RL agents exploit gaps in misspecified reward functions -- has been widely observed, but not yet systematically studied. To understand how reward hacking arises, we construct four RL environments with misspecified rewards. We investigate reward hacking as a function of agent capabilities: model capacity, action space resolution, observation space noise, and training time. More capable agents often exploit reward misspecifications, achieving higher proxy reward and lower true reward than less capable agents. Moreover, we find instances of phase transitions: capability thresholds at which the agent's behavior qualitatively shifts, leading to a sharp decrease in the true reward. Such phase transitions pose challenges to monitoring the safety of ML systems. To address this, we propose an anomaly detection task for aberrant policies and offer several baseline detectors.

Alexander Pan, Kush Bhatia, Jacob Steinhardt
arXiv:2201.03544 · cs.LG, cs.AI, stat.ML · submitted Jan 10, 2022 · updated Feb 14, 2022
abstract · pdf · html · ICLR 2022; 19 pages

add comment on HN