从不确定性到行动:学习引导 LLM 智能体
From Uncertainty to Action: Learning to Steer LLM Agents
Hanwen Li · Jinhao Duan · Guanhua Zhu · Junchi Lu · Bo Shen · Chenxi Yuan · et al.
中文摘要
引导一个 LLM 智能体意味着决定是否纠正它、在哪一步纠正以及使用何种机制。不确定性常被用来判断何时纠正智能体,但它能否指导这些决策尚不清晰。研究在每个非终止步骤上分别用四种机制引导智能体轨迹,并将每条延续运行至完成。由此得到的逐步结果表(SOT)包含来自三个基准和两个智能体的 1,864 条轨迹的约 82,000 条反事实延续。SOT 表明不确定性能够识别失败轨迹,但没有任何单一信号能可靠地定位引导起作用的步骤。因此提出 VoS(Value of Steering),一种轨迹级监控器,可离线或在线运行,从 SOT 中学习每一步的引导价值,并据此决定在何处介入。一个受伤害预算约束的触发器决定是否进行引导,限制 VoS 干扰成功轨迹的比例。在基准、智能体以及离线或在线使用的全部 12 种设置中,VoS 较未修改的执行平均提升 7.8 分;在 11 种设置中优于五种现有不确定性触发方法中最强者,平均提升 2.9 分。消融实验显示在已测量结果上训练与紧致的伤害预算均不可或缺。
关键要点
- 01问题:不确定性可识别失败轨迹,但无法可靠定位引导有效的步骤,缺乏统一的引导价值评估机制
- 02方法:构建包含约 82,000 条反事实延续的逐步结果表 SOT,并提出轨迹级监控器 VoS,学习每步引导价值后以伤害预算触发介入
- 03结果:在 12 种基准-智能体-在线/离线设置中较未修改执行平均提升 7.8 分,在 11 种中优于最强基线平均 2.9 分
- 04关键组件:在 SOT 已测量结果上训练与紧致伤害预算均为性能关键
- 05局限:仅覆盖三个基准与两个智能体,反事实空间局限于四种引导机制的组合
解读
尚无解读。
原始英文摘要
arXiv:2610.09115v1 Announce Type: cross Abstract: Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer agent trajectories separately at every non-terminal step with each of four mechanisms and run each continuation to completion. The resulting stepwise outcome table (SOT) holds about 82,000 counterfactual continuations of 1,864 trajectories from three benchmarks and two agents. It shows that uncertainty can identify failing trajectories, but that no single signal reliably locates the step at which steering helps. We therefore propose VoS (Value of Steering), a trajectory-level monitor, offline or online, that learns from SOT the value of steering at each step and decides where to steer by it. A harm-budgeted trigger decides whether to steer, limiting the fraction of successful trajectories that VoS disturbs. VoS improves on unmodified execution in all 12 settings of benchmark, agent, and offline or online use, by 7.8 points on average, and outperforms the strongest of five existing uncertainty-triggered methods in 11, by 2.9 points on average. Ablations show that training on measured outcomes and a tight harm budget are both essential.