跳到正文
返回论文列表
cs.LG提交于 已译

组合每位教师所学:基于教师相对偏移的多教师在线策略蒸馏

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

Hejian Sang · Zhengze Zhou · Shayan Mohajer Hamidi · Xiaomin Li · Rohit Jain · Alborz Geramifard

中文摘要

多教师在线策略蒸馏(MOPD)应用于两种场景。在同域组合(composition)中,多名教师对来自同一提示域的学生 rollout 各自评分,其信号合并为单一目标;在路由域蒸馏(routed-domain distillation)中,不同域的提示被分配给对应的专家教师。两种场景通常迁移每位教师的端点策略,这会把后训练带来的变化与教师基座模型继承来的偏好混在一起。论文提出 $\Delta$-MOPD,将每位教师的「教师减去基座」的 logit 偏移以学生冻结初始化为锚点重新对齐,并在固定教师选择的前提下,与两种场景中的端点监督进行对比。首先揭示阻碍端点迁移的机制:继承的基座牵引可能超过后训练偏移带来的效果。去除该基座成分后,教师项的范数比以及目标与学生之间的 KL 散度均下降。实验结果表明:当教师信号在同一个状态上组合时,偏移目标尤其有效。在三名组合教师下,$\Delta$-MOPD 比端点组合在 Math 上高出 4.11 分、在五项 benchmark 平均上高出 1.95 分;在两名教师下,两者准确率相当。在分阶段路由中,两种相位顺序下 $\Delta$-MOPD 均取得更高平均性能,并将观察到的顺序差距从 10.50 分缩小到 6.42 分。在交错路由(每次更新仅涉及一名教师)中,两种目标表现相近。分阶段结果为该收益可延伸至跨训练阶段累积的信号提供了支持证据。因此,目标构造是 MOPD 中独立于教师选择的设计维度,二者互补。

关键要点

  1. 01问题:多教师在线策略蒸馏迁移教师端点策略时,会把后训练变化与基座继承偏好混在一起,造成基座牵引过大、阻碍有效迁移
  2. 02方法:提出 $\Delta$-MOPD,以学生冻结初始化为锚点迁移「教师减去基座」的 logit 偏移,作为端点监督之外的新目标构造维度
  3. 03结果:三名教师组合下在 Math 领先 4.11 分、五项 benchmark 平均领先 1.95 分;两名教师下与端点监督持平
  4. 04结果:分阶段路由中在两种相位顺序下均更优,并将顺序差距从 10.50 分缩小到 6.42 分,交错路由下两者相当
  5. 05意义:目标构造是 MOPD 中独立于教师选择的设计轴,与教师选择互补,可推广至跨训练阶段累积的信号

解读

尚无解读。

原始英文摘要

arXiv:2610.10460v1 Announce Type: new Abstract: Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both settings usually transfer each teacher's endpoint policy, which mixes what post-training changed with preferences inherited from the teacher's base. We introduce $\Delta$-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed. We first expose the mechanism that impedes endpoint transfer: inherited base pull can exceed the post-training shift. Removing it reduces the teacher-term norm ratio and target--student KL. Across our experiments, the results suggest that shift targets are particularly useful when teacher signals are combined at a state. With three composed teachers, $\Delta$-MOPD exceeds endpoint composition by $4.11$ Math and $1.95$ five-benchmark points; with two, it matches endpoint accuracy. Under phased routing, it achieves higher mean performance in both phase orders and reduces the observed order gap from $10.50$ to $6.42$ points. Under interleaved routing, where each update involves one teacher, the two targets perform comparably. The phased results provide supporting evidence that the benefit may extend to signals accumulated across training phases. Target construction is thus an independent design axis in MOPD, complementary to teacher selection.

同方向论文 · cs.LG

查看全部 →