about
Search-R1: Training LLMs to Reason and Leverage Search Engines with RL (arxiv.org)
101 points by jonbaer on Apr 3, 2025 | hide | past | pdf | 12 comments on HN

In plain words: Instead of prompting a model to look things up, this trains it with rewards to write its own search queries mid-thought and fold results into its answers. On seven question-answering tests it beat the usual retrieve-then-answer approach by 41% with the larger of two models.

Abstract · Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Efficiently acquiring external knowledge and up-to-date information is essential for effective reasoning and text generation in large language models (LLMs). Prompting advanced LLMs with reasoning capabilities to use search engines during inference is often suboptimal, as the LLM might not fully possess the capability on how to interact optimally with the search engine. This paper introduces Search-R1, an extension of reinforcement learning (RL) for reasoning frameworks where the LLM learns to autonomously generate (multiple) search queries during step-by-step reasoning with real-time retrieval. Search-R1 optimizes LLM reasoning trajectories with multi-turn search interactions, leveraging retrieved token masking for stable RL training and a simple outcome-based reward function. Experiments on seven question-answering datasets show that Search-R1 improves performance by 41% (Qwen2.5-7B) and 20% (Qwen2.5-3B) over various RAG baselines under the same setting. This paper further provides empirical insights into RL optimization methods, LLM choices, and response length dynamics in retrieval-augmented reasoning. The code and model checkpoints are available at https://github.com/PeterGriffinJin/Search-R1.

Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, Jiawei Han
arXiv:2503.09516 · cs.CL, cs.AI, cs.IR · submitted Mar 12, 2025 · updated Aug 5, 2025
abstract · pdf · html · 31 pages

add comment on HN

This is the magical thing that happens when AI research happens in the open. Deepseek published their model and their methodology and then the nice people at the University of Illinois are able to build on it.

When OpenAI was launched this is what I thought it was going to be like. Something, something for the betterment of man kind.

I'm always surprised at how many LLM research papers are published on here, so despite OpenAI, I think it's absolutely happening.
Unfortunately the "open"AI effect is starting to show in other labs as well. DeepMind recently announced a min 6months delay in publishing their SotA research, to give them a market advantage. I get it, but it's sad that it's happening.

The good thing is that there are a lot of companies out there that want to make a name for themselves. Mistral started like that with Apache 2.0 models, now ds w/ MIT models, and so on. And if the past year is a good indicator, it seems that closed SotA to open close-to-SotA is 6-3 months. So that's good.

I also find interesting LeCun's take that "there is no closed source moat, or not for long". In a podcast he went into detail on this, saying that "people move companies, and people talk". If someone finds some secret sauce, the ideas will move around and other labs will catch up quickly. So there's some hope.

A couple of comments. What’s not that interesting here is that adding search to an LLM increases accuracy — this is known, and largely implemented via RAG or other search pipelines which then stuff information into the context.

What might be interesting here is that they are thinking about taxonomic tool use-cases, and exploring training and therefore optimizing the utilization of them.

This to me is a proof of concept — an interesting one, but just a proof of concept. You can see from their example search that the model over-relied on search; it didn’t need to re-search three times to get the answer.

A next step that I think would be useful would be updating the reward function to penalize search; pressing the model to use search when it needs to and not before. This to me is a likely framework going forward where MCP tool costing matters, and would be really useful to have in the next gen of tool calling LLMs.

In the case of search we’d hopefully get a really useful signal and outcome for times the model is unsure — it would call a friend, and get good info! And for times it’s sure, we’d have taught it not to waste reward on that.

This is pretty cool. I have a similar model that’s 8 days into training on msmarco.

So far I only have the “cold start” data posted, but I’m planning on posting a full distillation dataset.

https://huggingface.co/datasets/dleemiller/lm25

What kind of hardware setup would be needed to replicate the paper’s results?
I am training phi-4 (14B) using a single A6000. There’s some tricks you have to use to keep VRAM consumption down - mainly LoRA and quantization.

There’s a package called “unsloth” that integrates with huggingface’s TRL library that can help.

As far as I know, the idea behind Search-R1 stemmed from DeepRetrieval (search it on GitHub), though the latter has gained much less attention. Also, DeepRetrieval was trained using real search engines, not just BM25. If you check their training log, they got incredible performance (65% vs SOTA 25%) much earlier.
Leveraging reinforcement learning (RL) for LLMs is a fascinating evolution in search technology. The potential for improving search engines to reason intelligently and process data in real-time could revolutionize the entire industry.
Can someone ELI5 how reinforcement learning works with transformer based architecture?
I wonder if Perplexity uses similar methods under the hood or if it is a completely different approach.
I feel like most of these services simply take your prompt and ask a model for search queries regarding that prompt. Then add the resulting pages into the context.