about
Can reinforcement learning for LLMs scale beyond math and coding tasks? Probably (arxiv.org)
6 points by GabrielBianconi on Apr 8, 2025 | hide | past | pdf | 4 comments on HN

In plain words: They extend training that rewards good answers beyond math and coding to medicine, chemistry, and psychology, where answers are free-form: a small 7-billion-parameter model grades answers and uses that score instead of a right/wrong mark. This beat open models up to 72 billion parameters.

Abstract · Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reinforcement learning with verifiable rewards (RLVR) has demonstrated significant success in enhancing mathematical reasoning and coding performance of large language models (LLMs), especially when structured reference answers are accessible for verification. However, its extension to broader, less structured domains remains unexplored. In this work, we investigate the effectiveness and scalability of RLVR across diverse real-world domains including medicine, chemistry, psychology, economics, and education, where structured reference answers are typically unavailable. We reveal that binary verification judgments on broad-domain tasks exhibit high consistency across various LLMs provided expert-written reference answers exist. Motivated by this finding, we utilize a generative scoring technique that yields soft, model-based reward signals to overcome limitations posed by binary verifications, especially in free-form, unstructured answer scenarios. We further demonstrate the feasibility of training cross-domain generative reward models using relatively small (7B) LLMs without the need for extensive domain-specific annotation. Through comprehensive experiments, our RLVR framework establishes clear performance gains, significantly outperforming state-of-the-art open-source aligned models such as Qwen2.5-72B and DeepSeek-R1-Distill-Qwen-32B across domains in free-form settings. Our approach notably enhances the robustness, flexibility, and scalability of RLVR, representing a substantial step towards practical reinforcement learning applications in complex, noisy-label scenarios.

Yi Su, Dian Yu, Linfeng Song, Juntao Li, Haitao Mi, Zhaopeng Tu, Min Zhang, Dong Yu
arXiv:2503.23829 · cs.CL · submitted Mar 31, 2025 · updated Apr 1, 2025
abstract · pdf · html

add comment on HN

From the DeepSeek paper, they did try but found that the model would learn to cheat the judge. It doesn't seem to be impossible, but probably a serious challenge.

> We do not apply the outcome or process neural reward model in developing DeepSeek-R1-Zero, because we find that the neural reward model may suffer from reward hacking in the large-scale reinforcement learning process, and retraining the reward model needs additional training resources and it complicates the whole training pipeline.

To me, this is one of the most frustrating parts of this type of ML. If we could actually track the steps taken in the LLM, it would be trivial for the judge to evaluate the output of each intermediate and detect when reward hacking is taking place.

I wonder if there's any alternative other than trying to build the perfect judge for every single test case.

Recent paper from Anthropic on the fact that the reasoning output does not necessarily reflect what the LLM is "thinking".

https://arxiv.org/abs/2305.04388

That being said, your idea is not unreasonable. The way DeepSeek phrased it, it just sounds like implementing such solutions might be a hassle greatly increasing complexity, and they were just focused on making an RL baseline work at scale.

I was actually thinking of that paper when I wrote that comment, hence the frustration that we don't actually know the intermediates.

Still, perhaps the stepped output we get may hint at that kind of "cheating" and can be used in reinforcement... or perhaps that kind of reinforcement will just make the LLMs better at cheating. The problem is definitely a lot more complex than the trivial way I referenced it, at least.