哪一次 rollout 教会了它?BehaviorTrace 与在线强化学习中训练数据归因的局限
Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL
Amit Nautiyal
中文摘要
当强化学习教会语言模型一种新行为时,能否找到教会该行为的训练 rollout?当某种归因方法声称能找出时,又如何确认结论真实可靠?研究在 GRPO 在线强化学习微调场景下同时考察这两个问题,并使用具有已知成因的植入行为进行验证。BehaviorTrace 是一个开源评测框架,包含全梯度草图(full-gradient sketching)、植入行为设置,以及对梯度幅度、流利度、上升空间(headroom)、种子与生成采样变异性的控制。在 Qwen2.5-1.5B 上使用三个种子实验后发现,大量表观归因信号来自混淆因素。仅按梯度大小对训练步排序、不依赖任何行为目标的对照方法,就达到了 4.2 到 4.5 倍随机水平,并在三个种子中的两个上达到或超过最佳有目标估计器。在饱和检查点,模型流利度对行为标签的预测能力不弱于所有参与比较的梯度方法。一旦控制流利度,每次 rollout 的结果会因种子和生成采样不同而变化,因此单次运行无法得出确定结论。有一个信号在三个种子上都成立:触发词(trigger tokens)处的梯度与行为实际出现处构建的目标对齐。研究据此提出一份强化学习归因评估清单。测试中既包含现有估计器,包括 GAS(重归一化 TracInCP)和 TRAK 风格估计器,但未提出新方法。
关键要点
- 01问题:在线强化学习后语言模型习得新行为,但缺乏可靠手段定位具体是哪次 rollout 教会了该行为,且归因方法的有效性难以验证。
- 02方法:构建 BehaviorTrace 开源评测框架,结合全梯度草图、植入行为设置,控制梯度幅度、流利度、上升空间、种子与生成采样变异性。
- 03结果:纯梯度大小排序的对照方法达到 4.2–4.5 倍随机水平,在三个种子中的两个上不弱于有目标估计器,说明大量归因信号来自混淆。
- 04局限:控制流利度后单次 rollout 结果随种子和生成采样变化,唯一在三个种子上稳定成立的信号是触发词梯度与行为实际出现处的目标对齐。
- 05贡献:提出强化学习归因评估清单,评测 GAS(重归一化 TracInCP)与 TRAK 风格估计器,但未提出新归因方法。
解读
尚无解读。
原始英文摘要
arXiv:2610.10422v1 Announce Type: new Abstract: When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it? And when an attribution method says it can, how do we know the answer is real? We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause. We release BehaviorTrace, an open evaluation harness that combines full-gradient sketching, the planted-behavior setup, and controls for gradient magnitude, fluency, headroom, and variation across seeds and generation draws. Across three seeds on Qwen2.5-1.5B, much of the apparent attribution signal comes from confounds. A control that ranks training steps by gradient size alone, with no behavior target, reaches 4.2 to 4.5 times chance and matches or beats the best targeted estimator on two of three seeds. At saturated checkpoints, model fluency predicts the behavior label at least as well as every gradient method we compared it with. Once fluency is controlled, the per-rollout results change from seed to seed and from one generation draw to the next, so a single run cannot settle the question. One signal does hold on all three seeds. The gradient of the trigger tokens aligns with a target built where the behavior actually occurs. We turn these findings into a checklist for evaluating attribution in RL. We test existing estimators, including GAS (renormalized TracInCP) and a TRAK-style estimator, and do not propose a new one.