about
Value-Based Deep RL Scales Predictably (arxiv.org)
68 points by bearseascape on Feb 8, 2025 | hide | past | pdf | 3 comments on HN

In plain words: Measuring how the ratio of practice updates to new experience shifts the data-versus-compute tradeoff lets small runs forecast what a bigger run needs. Unlike the usual guess-and-check scaling, these forecasts held for three value-based reinforcement learning algorithms at larger budgets.

Abstract

Scaling data and compute is critical to the success of modern ML. However, scaling demands predictability: we want methods to not only perform well with more compute or data, but also have their performance be predictable from small-scale runs, without running the large-scale experiment. In this paper, we show that value-based off-policy RL methods are predictable despite community lore regarding their pathological behavior. First, we show that data and compute requirements to attain a given performance level lie on a Pareto frontier, controlled by the updates-to-data (UTD) ratio. By estimating this frontier, we can predict this data requirement when given more compute, and this compute requirement when given more data. Second, we determine the optimal allocation of a total resource budget across data and compute for a given performance and use it to determine hyperparameters that maximize performance for a given budget. Third, this scaling is enabled by first estimating predictable relationships between hyperparameters, which is used to manage effects of overfitting and plasticity loss unique to RL. We validate our approach using three algorithms: SAC, BRO, and PQL on DeepMind Control, OpenAI gym, and IsaacGym, when extrapolating to higher levels of data, compute, budget, or performance.

Oleh Rybkin, Michal Nauman, Preston Fu, Charlie Snell, Pieter Abbeel, Sergey Levine, Aviral Kumar
arXiv:2502.04327 · cs.LG · submitted Feb 6, 2025 · updated Jul 25, 2025
abstract · pdf · html · ICML 2025

add comment on HN

My attempt at a summary: authors characterize the data-compute pareto front (aka how bitter is the lesson, exactly?)

For a different perspective, error vs compute, see

https://youtu.be/5eqRuVp65eY

and comments

(I particularly liked the one about string theorists rediscovering a fundamental theorem in GR decades too late-- rediscovering how to integrate happens in every field, it's nothing to be ashamed of :)

Skimmed little bits: "on policy" RL means the model has generated output and received feedback from sort of dynamic environment, which might not be scalable. Value-based off-policy means the model is trained with data that wasn't generated from the model itself exploring a dynamic environment. Instead it can be recordings. They then ask the question; how does that scale?
RL is unbelievably finicky, sensitive to hyperparameters, and hard to make work. If you are pursuing a research project and decide to use RL, you are making your life a lot more difficult and stressful.

It's exciting to see any progress in making accurate predictions about what settings will work for RL training. I hope that this research direction can be expanded in scope and that ultimately, people who want to do research in RL can become confident in their training recipes.

I am more excited about that, than about the dream of scaling compute per se.