about
Reinforcement Learning Under Moral Uncertainty (arxiv.org)
33 points by hardmaru on Jul 17, 2020 | hide | past | pdf | 6 comments on HN

In plain words: Agents are trained with several moral theories at once, weighting each by how plausible it seems, instead of being rewarded for following just one. In simple tests this cut down the extreme actions single-theory agents take, though balancing rewards that can't be compared stays tricky.

Abstract

An ambitious goal for machine learning is to create agents that behave ethically: The capacity to abide by human moral norms would greatly expand the context in which autonomous agents could be practically and safely deployed, e.g. fully autonomous vehicles will encounter charged moral decisions that complicate their deployment. While ethical agents could be trained by rewarding correct behavior under a specific moral theory (e.g. utilitarianism), there remains widespread disagreement about the nature of morality. Acknowledging such disagreement, recent work in moral philosophy proposes that ethical behavior requires acting under moral uncertainty, i.e. to take into account when acting that one's credence is split across several plausible ethical theories. This paper translates such insights to the field of reinforcement learning, proposes two training methods that realize different points among competing desiderata, and trains agents in simple environments to act under moral uncertainty. The results illustrate (1) how such uncertainty can help curb extreme behavior from commitment to single theories and (2) several technical complications arising from attempting to ground moral philosophy in RL (e.g. how can a principled trade-off between two competing but incomparable reward functions be reached). The aim is to catalyze progress towards morally-competent agents and highlight the potential of RL to contribute towards the computational grounding of moral philosophy.

Adrien Ecoffet, Joel Lehman
arXiv:2006.04734 · cs.AI · submitted Jun 8, 2020 · updated Jul 19, 2021
abstract · pdf · html · 28 pages, 18 figures; update adds discussion of a possible flaw of Nash voting, discussion of further possible research into MEC, as well as a few more references; updated to ICML version

add comment on HN

As I see it, morality is inherently underdefined and subjective. All our current attempts at creating rigorous moral frameworks lead to intuitively immoral behaviour under some circumstances. Combining them all together with some sort of blending function to avoid the weak points might avoid the kookiest of unintuitive behaviour but I don't think it'll solve the inherent fuzziness of actual morality.
Fortunately, it doesn't have to solve it to be useful. It just has to approximate human-quality morality. Humans are well known for being, shall we say, fuzzy in their morality at times.
> All our current attempts at creating rigorous moral frameworks lead to intuitively immoral behaviour under some circumstances.

Objectivism does not.

I'm not certain if that was a joke or not. Calling altruism immoral certainly qualifies as intuitively immoral behavior.

Even if one excises Ayn Rand's infamous politics and personal hypocrisies, and even everything which doesn't follow from claimed principles it isn't very workable. Rationality is a measure of sanity more than morality even if there may be some overlap in that a perfectly rational actor wouldn't display gratitutous cruelty.

Granted everything we have is also imperfect. Pure intuition can easily fall into nonsensical superstition and prejudices like Pythagorean hatred for beans and irrational number denial and rationality alone. There are plenty of "serial killer organ thief doctor utilitarianism" arguments but those first order calculations fail to consider the impact of it inevitably becoming a known thing in society. Being incentivized to shoot a doctor if you wind up alone with one is a larger net harm to society in addition to the obvious deterrant to being a lone traveler. Those silly exercises aside rationality alone is insufficient it lacks goal definition. One may rationally persue extinction of all life in the universe or maximizing human lifespan.

If a sub-goal is linked to another goal it can be found to be irrational if inconsistent, if painting your dog won't make your crush love you dog-painting isn't a rational goal. But if you want to just paint your dog for the sake of doing so (please don't) it can't be said to be less rational than another goal even if it is more or less obtainable.

The Stanford Encyclopedia of Philosophy recently (March 2020) published an entry on Computational Philosophy, a newly forming field that employs "the use of mechanized computational techniques to instantiate, extend, and amplify philosophical research."

https://plato.stanford.edu/entries/computational-philosophy

This paper is interesting enough, but the authors also released code for their environments and experiments used in the paper: https://github.com/uber-research/normative-uncertainty