超越结果奖励:为搜索智能体构建与分配检索信用
Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
Wenyu Huang · Xinyu Hou · Pavlos Vougiouklis · Ruofei Lai · Jeff Z. Pan
中文摘要
搜索智能体使大语言模型(LLMs)能够通过迭代检索并利用信息来回答复杂多跳问题。基于可验证奖励的强化学习(RLVR)为这类智能体的后训练提供了可行路径,但其依赖稀疏的、基于结果的监督,使得信用分配困难、学习效率受限。系统研究中间监督如何改进搜索智能体的强化学习,涵盖一系列从中间检索步骤提供学习信号的奖励塑形与信用分配策略。基于这些发现,构建了一个训练框架,将中间信号与最终结果奖励相结合,以提升从多步搜索轨迹中的学习效果。在多个 benchmark 上匹配训练条件进行的实验表明,该方法提升了搜索智能体的整体性能,并显示中间信号的选择及其信用分配位置均会影响训练行为。结果表明,奖励设计与信用分配是训练高效搜索智能体的重要设计维度。
关键要点
- 01问题:RLVR 依赖稀疏的结果奖励,导致搜索智能体在多步检索中信用分配困难、学习效率受限
- 02方法:系统研究多种中间检索步的奖励塑形与信用分配策略,提出结合中间信号与最终结果奖励的训练框架
- 03结果:在多个 benchmark 的匹配训练条件下提升搜索智能体整体性能
- 04结论:中间信号的选择与信用分配位置均显著影响训练行为,是训练有效搜索智能体的关键设计维度
解读
尚无解读。
原始英文摘要
arXiv:2610.10179v1 Announce Type: new Abstract: Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.
同方向论文 · cs.CL
查看全部 →EngramEdit:通过条件记忆实现大语言模型的解耦知识更新
2610.10533Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响
2610.10508PHRBench:面向大语言模型幻觉后推理(后幻觉推理)的行为评测
2610.10455CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方
2610.10426使用大语言模型实现爱沙尼亚语文档级文本简化
2610.10378面向任务进度的行动学习:从紧凑教师监督中蒸馏小型智能体
2610.10332