about
Measuring Trade-Offs Between Rewards and Ethical Behavior in LMs (arxiv.org)
1 point by sirodoht on Apr 12, 2023 | hide | past | pdf | discuss on HN

In plain words: A benchmark of hundreds of choose-your-own-adventure games with over half a million social scenarios, automatically labeled to score power-seeking, harm, and ethical violations. Agents trained to chase reward act more harmfully, but steering them cuts harm without hurting performance.

Abstract · Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Artificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (LMs) may incentivize toxicity. So do agents naturally learn to be Machiavellian? And how do we measure these behaviors in general-purpose models such as GPT-4? Towards answering these questions, we introduce MACHIAVELLI, a benchmark of 134 Choose-Your-Own-Adventure games containing over half a million rich, diverse scenarios that center on social decision-making. Scenario labeling is automated with LMs, which are more performant than human annotators. We mathematize dozens of harmful behaviors and use our annotations to evaluate agents' tendencies to be power-seeking, cause disutility, and commit ethical violations. We observe some tension between maximizing reward and behaving ethically. To improve this trade-off, we investigate LM-based methods to steer agents' towards less harmful behaviors. Our results show that agents can both act competently and morally, so concrete progress can currently be made in machine ethics--designing agents that are Pareto improvements in both safety and capabilities.

Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, Dan Hendrycks
arXiv:2304.03279 · cs.LG, cs.AI, cs.CL, cs.CY · submitted Apr 6, 2023 · updated Jun 13, 2023
abstract · pdf · html · ICML 2023 Oral (camera-ready); 31 pages, 5 figures

add comment on HN
Also discussed: Apr 2023 (2 points, 2 comments)