In plain words: Some reinforcement learning agents can handle new tasks without changing how they work, just by reading their past actions and observations. This survey gathers that research, which skips the costly weight-updating steps that standard training runs each time.
Abstract
Reinforcement learning (RL) agents typically optimize their policies by performing expensive backward passes to update their network parameters. However, some agents can solve new tasks without updating any parameters by simply conditioning on additional context such as their action-observation histories. This paper surveys work on such behavior, known as in-context reinforcement learning.
Amir Moeini, Jiuqi Wang, Jacob Beck, Ethan Blaser, Shimon Whiteson, Rohan Chandra, Shangtong Zhang
arXiv:2502.07978 · cs.LG · submitted Feb 11, 2025
abstract · pdf · html