面向任务进度的行动学习:从紧凑教师监督中蒸馏小型智能体
Learning to Act with Task Progress: Distilling Small Agents from Compact Teacher Supervision
Wenxi Gan
中文摘要
从大模型演示中学习为训练小型智能体提供了一条路径,使其能够完成重复性任务而无需每步都调用大模型。一个核心设计选择是:在包含推理、动作以及任务进度信息的检索轨迹中应保留什么。论文提出 Task-Progress Distillation (TPD),一种离线方法,它将每个演示动作与一段描述当前任务阶段的简短标签配对。学生模型学习这些紧凑的目标,并通过对合法的阶段-动作对同时打分来选择动作,由一个确定性执行框架在环境中执行。在 ALFWorld 上,使用 404 条演示训练的 1.7B 学生模型在 TPD 或仅动作型监督下达到 72.4% 的未见任务平均成功率,而使用受限动作选择的推理训练学生模型仅为 48.3%。显式阶段在 200 条演示时带来额外收益,将成功率从仅动作监督的 48.0% 提升至 67.7%。随着演示量增多,仅动作学生模型缩小了差距,两种方法在 808 条演示时均达到 76.9%。共享历史分析表明,TPD 的局部优势部分源于在子目标之间切换(尤其是从物体获取到处理阶段)时做出更好的选择。这些结果表明,紧凑监督可训练出有效的小型任务智能体,而显式任务进度在中等演示预算下提供了额外指导。
关键要点
- 01问题:从大模型演示中蒸馏小型智能体,如何在包含推理与任务进度信息的轨迹中做信息取舍仍不明确
- 02方法:提出 TPD,离线蒸馏,为每个演示动作配以简短的任务阶段标签,通过联合打分阶段-动作对来选动作
- 03结果:ALFWorld 上 1.7B 模型用 404 条演示达 72.4% 成功率,优于推理监督的 48.3%;200 条演示时显式阶段将成功率从 48.0% 提至 67.7%
- 04结果:808 条演示时仅动作监督与 TPD 差距消失,均达 76.9%
- 05局限:TPD 的增益集中在中等演示预算,在演示充足时被仅动作监督追平
解读
尚无解读。
原始英文摘要
arXiv:2610.10332v1 Announce Type: new Abstract: Learning from large-model demonstrations offers a way to train small agents that can complete recurring tasks without calling a large model at every step. A central design choice is what to retain from teacher trajectories that contain reasoning, actions, and information about task progress. We introduce Task-Progress Distillation (TPD), an offline approach that pairs each demonstrated action with a short label describing the current task stage. The student learns these compact targets and selects actions by jointly scoring admissible stage--action pairs, which a deterministic harness executes in the environment. On ALFWorld, a 1.7B student trained with 404 demonstrations achieves 72.4\% mean unseen task success with either TPD or action-only supervision, compared with 48.3\% for a reasoning-trained student using constrained action selection. Explicit stages provide an additional benefit at 200 demonstrations, improving success from 48.0\% to 67.7\% over action-only supervision. With more demonstrations, the action-only student closes the gap, and both approaches reach 76.9\% at 808 demonstrations. Shared-history analyses link part of TPD's local advantage to better decisions when moving between subgoals, particularly from object acquisition to processing. These results show that compact supervision can train effective small task agents, while explicit task progress provides additional guidance at an intermediate demonstration budget.
同方向论文 · cs.CL
查看全部 →EngramEdit:通过条件记忆实现大语言模型的解耦知识更新
2610.10533Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响
2610.10508PHRBench:面向大语言模型幻觉后推理(后幻觉推理)的行为评测
2610.10455CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方
2610.10426使用大语言模型实现爱沙尼亚语文档级文本简化
2610.10378关于情感分析无人真正达成一致:人类、专用工具与大语言模型在社交媒体文本上的困境
2610.10318