跳到正文
返回论文列表
cs.AI提交于 已译

面向语言反馈学习的约束树探索

Constraint Tree Exploration for Learning from Language Feedback

Shaoang Li · Daniel R. Jiang · Jian Li

中文摘要

交互式学习中的自然语言反馈通常会指出违反的需求,从而解释某个动作为何失败。错误解读反馈会导致其排除本应成立的解。通过将用户意图建模为动作空间上的潜在约束,并把从语言反馈学习形式化为对可行域的纯探索问题,对该设定展开研究。提出 TRACE 算法:将候选约束组织成一棵树,通过生成满足待测约束的动作来检验每个候选细化。仅当多次测试产生的反馈与之不矛盾时,才采纳该细化。区分同一反馈的两种用途:(i) falsification,用于检测与当前所测约束集的矛盾;(ii) identification,可额外指出被违反的约束。证明 TRACE-Falsification 具有关于候选类规模 H 的高概率覆盖界;在可靠 identification 下,TRACE-Identification 可将该依赖替换为 K/p_ext,其中 K 为潜在约束数量,p_ext 为从信息性反馈中提取缺失真实约束的概率下界。在 6 个语言反馈任务上评估 TRACE。在 RecMovie 任务中,TRACE-Identification 在评估输出上限分别为 20 和 60 时,最终输出成功率分别达到 73% 和 86%,而在相同反馈与输出上限下,所评测的提示基线方法至多仅达到 42% 和 48%。受控的身份污染实验进一步表明,在 falsification 检测器可靠时,该方法相比直接累加策略具有更强的鲁棒性。

关键要点

  1. 01问题:自然语言反馈易被曲解,导致智能体排除本应成立的解,需要将用户意图建模为动作空间上的潜在约束
  2. 02方法:提出 TRACE 算法,把候选约束组织成约束树并通过反复生成满足动作进行检验,区分 falsification 与 identification 两种反馈利用方式
  3. 03理论:TRACE-Falsification 给出关于候选类规模 H 的高概率覆盖界;可靠 identification 下,TRACE-Identification 将其替换为 K/p_ext
  4. 04结果:在 RecMovie 任务上,TRACE-Identification 在输出上限 20 和 60 时分别达到 73% 和 86% 成功率,显著优于提示基线
  5. 05局限:identification 的收益依赖于可靠的约束提取,身份污染实验中在 falsification 检测器失效时鲁棒性下降

解读

尚无解读。

原始英文摘要

arXiv:2610.09107v1 Announce Type: cross Abstract: Natural-language feedback in interactive learning often explains why an action failed by pointing to violated requirements. Misinterpreting this feedback can lead an agent to rule out valid solutions. We study this setting by modeling user intent as latent constraints over an action space and formulating learning from language feedback as pure exploration over feasible regions. We introduce TRACE, an algorithm that organizes candidate constraints in a tree and tests each proposed refinement by generating actions that satisfy it. TRACE commits to the refinement only if the resulting feedback does not contradict it over repeated tests. We distinguish two ways of using the same feedback: (i) falsification, which detects contradictions to the constraint set currently being tested, and (ii) identification, which may additionally name a violated constraint. We prove high-probability coverage bounds with dependence on the candidate class size $H$ for TRACE-Falsification. With reliable identification, TRACE-Identification can replace this dependence by $K/p_{\mathrm{ext}}$, where $K$ is the number of latent constraints and $p_{\mathrm{ext}}$ lower-bounds the probability of extracting a missing true constraint from informative feedback. We evaluate TRACE across six language-feedback tasks. On RecMovie, TRACE-Identification achieves 73% and 86% final-output success under caps of 20 and 60 evaluated outputs, compared with at most 42% and 48% for the evaluated prompting baselines given the same feedback and output caps. Controlled identity-corruption experiments further show greater robustness than direct accumulation when the falsification detector remains reliable.

同方向论文 · cs.AI

查看全部 →