跳到正文
返回论文列表
cs.CL提交于 已译

CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

Jixuan Chen · Jiaxin Zhang · Qinyuan Ye · Yada Pruksachatkun · Haoxiang Zhang · Jingming Zhuo · et al.

中文摘要

终端 Agent 能力同时取决于模型权重与运行时 harness(负责格式化提示、绑定工具、处理错误恢复)。现有 harness-model 协同演化方法同时改进两者,但往往把 harness 搜索阶段产生的轨迹视作无差别的回放缓冲区,忽视了轨迹对模型训练的价值依赖于其生成时所使用的 harness。为系统分析这一接口,论文建立了交替协同演化框架,通过逐组件的提升决策解耦 harness 搜索与策略训练。在此框架下提出 CoTrace,一种 harness 感知的数据配方,显式管理轨迹路由、来源匹配与课程刷新。在 CoTrace 中,反复出现的执行失败引导 harness 合成,策略训练则严格基于已验证的 rollout:监督微调(SFT)使用与所采用运行时匹配的 rollout,强化学习(RL)使用全新的在线交互。在 Tmax 提升划分上,CoTrace 将 Qwen3.5-9B 的已解决问题数从 78 提升至 88(SFT),在线强化学习变体进一步达到 90。具体而言,一个紧凑的 harness 匹配语料库以远低于跨兄弟 harness 池化的大规模语料库的算力,即可带来稳定的模型增益。此外,在 Terminal-Bench 2.1 和 SWE-bench Lite 上的评测表明,分布外迁移本质上依赖 harness 兼容性,保持训练与评估运行时的一致性可避免在外部脚手架下出现的过程性执行崩溃。

关键要点

  1. 01现有 harness-model 协同演化把搜索轨迹当作无差别回放缓冲,忽视轨迹价值依赖生成时所用 harness
  2. 02提出交替协同演化框架,通过逐组件提升决策解耦 harness 搜索与策略训练
  3. 03设计 harness 感知的数据配方 CoTrace,显式管理轨迹路由、来源匹配与课程刷新:SFT 用匹配的已验证 rollout,RL 用全新在线交互
  4. 04在 Tmax 提升划分上,CoTrace 将 Qwen3.5-9B 已解决问题数从 78 提升至 88(SFT)与 90(在线 RL)
  5. 05紧凑的 harness 匹配语料库以更低算力取得稳定增益;Terminal-Bench 2.1 与 SWE-bench Lite 上,训练与评估 harness 不一致会导致过程性执行崩溃

解读

尚无解读。

原始英文摘要

arXiv:2610.10426v1 Announce Type: new Abstract: Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This practice overlooks that a trajectory's value for model training depends on the harness under which it was generated. To systematically analyze this interface, we establish an alternating co-evolution framework that decouples harness search and policy training through component-wise promotion decisions. Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh. Under CoTrace, recurring execution failures guide harness synthesis, while policy training is strictly conditioned on verified rollouts matched to the adopted runtime for supervised fine-tuning (SFT) or fresh online interactions for reinforcement learning (RL). On the Tmax promotion split, CoTrace advances Qwen3.5-9B from 78 to 88 solved tasks under supervised fine-tuning while an online reinforcement variant reaches 90. Specifically, a compact harness-matched corpus produces steady model gains at substantially lower compute than much larger corpora pooled across sibling harnesses. Furthermore, evaluations on Terminal-Bench 2.1 and SWE-bench Lite show that out-of-distribution transfer depends fundamentally on harness compatibility, where maintaining consistency between training and evaluation runtimes prevents procedural execution breakdowns observed under foreign scaffolds.

同方向论文 · cs.CL

查看全部 →