跳到正文
返回论文列表
cs.CL提交于 已译

超越结果奖励:为搜索智能体构建与分配检索信用

Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents

Wenyu Huang · Xinyu Hou · Pavlos Vougiouklis · Ruofei Lai · Jeff Z. Pan

中文摘要

搜索智能体使大语言模型(LLMs)能够通过迭代检索并利用信息来回答复杂多跳问题。基于可验证奖励的强化学习(RLVR)为这类智能体的后训练提供了可行路径,但其依赖稀疏的、基于结果的监督,使得信用分配困难、学习效率受限。系统研究中间监督如何改进搜索智能体的强化学习,涵盖一系列从中间检索步骤提供学习信号的奖励塑形与信用分配策略。基于这些发现,构建了一个训练框架,将中间信号与最终结果奖励相结合,以提升从多步搜索轨迹中的学习效果。在多个 benchmark 上匹配训练条件进行的实验表明,该方法提升了搜索智能体的整体性能,并显示中间信号的选择及其信用分配位置均会影响训练行为。结果表明,奖励设计与信用分配是训练高效搜索智能体的重要设计维度。

关键要点

  1. 01问题:RLVR 依赖稀疏的结果奖励,导致搜索智能体在多步检索中信用分配困难、学习效率受限
  2. 02方法:系统研究多种中间检索步的奖励塑形与信用分配策略,提出结合中间信号与最终结果奖励的训练框架
  3. 03结果:在多个 benchmark 的匹配训练条件下提升搜索智能体整体性能
  4. 04结论:中间信号的选择与信用分配位置均显著影响训练行为,是训练有效搜索智能体的关键设计维度

解读

尚无解读。

原始英文摘要

arXiv:2610.10179v1 Announce Type: new Abstract: Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.

同方向论文 · cs.CL

查看全部 →