about
Neural Networks with Motivation (arxiv.org)
2 points by signa11 on Sep 6, 2024 | hide | past | pdf | discuss on HN

In plain words: A motivation signal that scales expected rewards lets a learning network switch behavior instantly, without retraining its connections. In a conditioning task, its neurons split into two opposite groups, matching real brain cells that track positive and negative rewards.

Abstract · Neural networks with motivation

How can animals behave effectively in conditions involving different motivational contexts? Here, we propose how reinforcement learning neural networks can learn optimal behavior for dynamically changing motivational salience vectors. First, we show that Q-learning neural networks with motivation can navigate in environment with dynamic rewards. Second, we show that such networks can learn complex behaviors simultaneously directed towards several goals distributed in an environment. Finally, we show that in Pavlovian conditioning task, the responses of the neurons in our model resemble the firing patterns of neurons in the ventral pallidum (VP), a basal ganglia structure involved in motivated behaviors. We show that, similarly to real neurons, recurrent networks with motivation are composed of two oppositely-tuned classes of neurons, responding to positive and negative rewards. Our model generates predictions for the VP connectivity. We conclude that networks with motivation can rapidly adapt their behavior to varying conditions without changes in synaptic strength when expected reward is modulated by motivation. Such networks may also provide a mechanism for how hierarchical reinforcement learning is implemented in the brain.

Sergey A. Shuvaev, Ngoc B. Tran, Marcus Stephenson-Jones, Bo Li, Alexei A. Koulakov
arXiv:1906.09528 · q-bio.NC, cs.LG · submitted Jun 23, 2019 · updated Nov 19, 2019
abstract · pdf · html · Added the Methods section

add comment on HN