跳到正文
返回论文列表
cs.CL提交于 已译

基于任务结果为大语言模型智能体训练顾问(Advisors)

Training Advisors for LLM Agents from Task Outcomes

Sergei Polezhaev · Barys Liskavets · Ori Press · Alexander Golubev

中文摘要

大语言模型(LLM)智能体通过推理与工具调用与环境观察的交替执行来完成多步任务。已有研究表明,自然语言反馈可帮助这些智能体在任务过程中修正决策。论文提出 Caddie,一种训练评论器(critic)的方法,使其在智能体执行任务过程中提供自然语言分析与建议。与依赖逐步标注或参考评论的方法不同,Caddie 从智能体在接受评论器反馈后是否最终成功这一信号中学习。该方法通过强化学习优化评论器,同时保持基座模型冻结。论文使用单一基座模型在多跳问答任务上训练评论器,所得的 Qwen3-4B 评论器可改善四个不同规模与架构的基座模型的成功率,其中三个未在评论器训练中使用。在 MuSiQue 基准上,训练后的评论器将 Qwen3-4B 的成功率提升超过 25 个百分点,超越未使用评论器的 Kimi K3。同一评论器还在 τ³ 和 DeepDive 等域外交互式基准上带来增益,无需额外训练。结果表明,智能体能够在推理时决定何时向评论器寻求帮助,且基于结果的评论器优化训练可产生跨基座模型和任务域迁移的指导。

关键要点

  1. 01问题:现有依赖逐步标注或参考评论训练评论器的方法成本高、难以扩展,缺乏从任务最终结果学习的有效方式
  2. 02方法:提出 Caddie,通过强化学习从智能体最终是否成功这一结果信号训练评论器,基座模型保持冻结
  3. 03结果:在 MuSiQue 基准上将 Qwen3-4B 成功率提升超过 25 个百分点,超越 Kimi K3,且对三个未参与训练的基座模型均有效
  4. 04泛化性:同一评论器在 τ³ 和 DeepDive 等域外交互式基准上无需额外训练即可带来性能提升
  5. 05机制:智能体可在推理时自主决定何时向评论器寻求帮助,实现灵活的人机协作式决策

解读

尚无解读。

原始英文摘要

arXiv:2610.09858v1 Announce Type: new Abstract: Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B's success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including $\tau^3$ and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.

同方向论文 · cs.CL

查看全部 →