跳到正文
返回论文列表
cs.LG提交于 已译

好老师遇学生于其所在之处:策略内(on-policy)学习与教学的联合训练

A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching

Randy Ardywibowo · Arnav Dalal · Jiantao Jiao

中文摘要

基于结果奖励的强化学习(RL)存在监督稀疏的问题,在困难且长视野的任务中尤为突出,因为成功轨迹稀少、生成代价高。策略内蒸馏(OPD)提供了一种颇具吸引力的替代方案,由更强的教师沿学生自身的生成过程提供密集的逐 token 监督。自蒸馏(self-distillation)方法则通过让同一策略在特权信息(privileged information)条件下充当自身教师,从而省去了单独的教师模型。然而,仅凭特权条件并不能保证所得蒸馏更新能改善学生表现:特权信息可能让教师通过学生无法获得的捷径完成任务,产生与学生当前行为不匹配的监督;即便更高性能的教师,其指导也可能反而降低学生表现。为此,本文分析了特权教师的选择如何影响学生更新,推导出教师局部蒸馏更新成为学生奖励梯度正倍数的充要条件,并指出教师不仅要在任务上表现良好,还应提供与学生当前能力相适应的指导。基于此,本文提出一种实用的教师训练替代目标,将结果奖励与面向学生的逐 token KL 散度正则相结合,并据此设计联合策略内学习与教学(JOLT)方法,由单一策略同时扮演两个角色:以 KL 正则目标训练的特权教师,以及以密集策略内蒸馏训练的无特权学生。在数学推理、编程、工具调用和命令行使用等任务上,JOLT 提升了训练效率与性能,叠加学生奖励后还能进一步带来增益。

关键要点

  1. 01问题:基于结果奖励的强化学习监督稀疏;特权信息自蒸馏中,教师可能沿学生无法获得的捷径给出与学生当前能力不匹配的监督,导致更高性能教师反而降低学生表现。
  2. 02方法:推导出教师蒸馏更新与学生奖励梯度方向一致的充要条件,据此构造结合结果奖励与面向学生 KL 正则的教师训练目标,提出联合策略内学习与教学(JOLT),由同一策略同时担任特权教师与无特权学生。
  3. 03结果:在数学推理、编程、工具调用与命令行使用任务上,JOLT 提升训练效率与最终性能,额外加入学生奖励后获得进一步增益。
  4. 04局限:摘要未给出具体模型规模、参数量与 benchmark 数值,亦未讨论与现有自蒸馏/OPD 方法的直接对比细节。
  5. 05适用性:该教师训练准则依赖特权信息的存在,需在可获得特权上下文的强化学习场景中部署。

解读

尚无解读。

原始英文摘要

arXiv:2610.10447v1 Announce Type: new Abstract: Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting distillation update improves the student. Indeed, privileged information can lead the teacher to solve tasks through shortcuts unavailable to the student, producing supervision poorly matched to the student's current behavior. Consequently, even a higher-performing teacher can provide guidance that degrades student performance. To address this, we analyze how the choice of privileged teacher affects the student's update. We derive a necessary and sufficient condition for the teacher's local distillation update to be a positive multiple of the student's reward gradient. Our analysis suggests that the teacher should not only perform well on the task, but also provide guidance suited to the student's current capabilities. This characterization motivates a practical teacher-training surrogate that combines outcome rewards with token-level Kullback-Leibler (KL) regularization toward the student. Based on this result, we propose Joint On-Policy Learning and Teaching (JOLT), which jointly trains a single policy in two roles: a privileged teacher using a KL-regularized objective, and an unprivileged student using dense on-policy distillation. Across mathematical reasoning, coding, tool use, and terminal use, JOLT improves training efficiency and performance, with further gains from student rewards.

同方向论文 · cs.LG

查看全部 →