---
title: "R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning"
canonical_url: "https://www.modelscope.cn/papers/124030"
md_url: "https://www.modelscope.cn/papers/124030.md"
arxiv_id: 2503.05592
published: 2025-03-07
last_updated: 2025-03-07
authors:
  - "Huatong Song"
  - "Jinhao Jiang"
  - "Yingqian Min"
  - "Jie Chen"
  - "Zhipeng Chen"
  - "Wayne Xin Zhao"
  - "Lei Fang"
  - "Ji-Rong Wen"
model_name: R1-Searcher
model_developer: "中国人民大学高瓴人工智能学院"
domain:
  - "自然语言处理"
  - "深度学习"
  - "机器学习"
type:
  - "自然语言处理"
  - "深度学习"
  - "机器学习"
  - "Artificial Intelligence (cs.AI)"
  - "Computation and Language (cs.CL)"
  - "Information Retrieval (cs.IR)"
arxiv_url: "https://arxiv.org/abs/2503.05592"
pdf_url: "https://arxiv.org/pdf/2503.05592.pdf"
---

# R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

> Existing Large Reasoning Models (LRMs) have shown the potential of reinforcement learning (RL) to enhance the complex reasoning capabilities of Large Language Models~(LLMs). While they achieve remarkable performance on challenging tasks such as mathematics…

「R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning」是 ModelScope 魔搭社区收录的论文，arXiv 2503.05592，作者为 Huatong Song, Jinhao Jiang, Yingqian Min et al.，发表于 2025-03-07，属于 自然语言处理、深度学习、机器学习 领域。

- **ArXiv**: 2503.05592
- **Published**: 2025-03-07
- **Authors**: Huatong Song, Jinhao Jiang, Yingqian Min, Jie Chen, Zhipeng Chen, Wayne Xin Zhao, Lei Fang, Ji-Rong Wen
- **Model**: R1-Searcher
- **Developer**: 中国人民大学高瓴人工智能学院
- **Domain**: 自然语言处理, 深度学习, 机器学习
- **ArXiv URL**: https://arxiv.org/abs/2503.05592
- **PDF**: https://arxiv.org/pdf/2503.05592.pdf

Source: https://www.modelscope.cn/papers/124030

---

> R1-Searcher：通过强化学习激励大语言模型的搜索能力

## 摘要

现有的大型推理模型（LRMs）通过强化学习（RL）增强了大语言模型（LLMs）的复杂推理能力，但在处理时间敏感或知识密集型问题时，由于依赖内部知识，容易产生不准确性和幻觉。为了解决这一问题，本文提出了R1-Searcher，一种基于两阶段结果驱动的RL方法，旨在增强LLMs的搜索能力。该方法允许LLMs在推理过程中自主调用外部搜索系统以获取额外知识。R1-Searcher框架完全依赖于RL，无需过程奖励或蒸馏来启动冷启动。实验结果表明，R1-Searcher显著优于之前的强RAG方法，甚至超过了封闭源代码的GPT-4o-mini。具体而言，R1-Searcher在HotpotQA和2WikiMultiHopQA上的性能分别提升了48.22%和21.72%，并且在未见过的Bamboogle数据集上也表现出色，相比Search-o1提高了11.4%。R1-Searcher的核心贡献在于：1）引入了两阶段RL框架，使模型能够自主进行检索；2）在四个多跳问答基准测试中，R1-Searcher始终显著超越现有RAG方法；3）仅使用RL进行训练，无需任何蒸馏或冷启动，并展示了强大的泛化能力。

## Abstract

Existing Large Reasoning Models (LRMs) have shown the potential of reinforcement learning (RL) to enhance the complex reasoning capabilities of Large Language Models~(LLMs). While they achieve remarkable performance on challenging tasks such as mathematics and coding, they often rely on their internal knowledge to solve problems, which can be inadequate for time-sensitive or knowledge-intensive questions, leading to inaccuracies and hallucinations. To address this, we propose \textbf{R1-Searcher}, a novel two-stage outcome-based RL approach designed to enhance the search capabilities of LLMs. This method allows LLMs to autonomously invoke external search systems to access additional knowledge during the reasoning process. Our framework relies exclusively on RL, without requiring process rewards or distillation for a cold start. % effectively generalizing to out-of-domain datasets and supporting both Base and Instruct models. Our experiments demonstrate that our method significantly outperforms previous strong RAG methods, even when compared to the closed-source GPT-4o-mini.
