在 RLVR 中将探索与优化解耦
Decoupling Exploration from Optimization in RLVR
Saif Punjwani · Micah Goldblum
中文摘要
现代语言模型在已完成训练的 checkpoint 之上,采用带可验证奖励的强化学习(RLVR)继续训练。RLVR 的一项关键潜力在于发现新的推理策略,模型原则上可以采样出其训练数据中不存在的新颖思路。然而实践中,在 RLVR 中加入强新颖性激励的效果有限,且可能损害模型质量。由于可验证奖励只监督模型知识与行为的一小部分,这类退化难以恢复。为此,本文提出一个将探索与优化解耦的框架,称为探索-蒸馏(Exploration-Distillation,ExpDis):用一个或多个探索者策略以含新颖性奖励进行训练,对其轨迹的正确性与质量进行过滤,再蒸馏到独立的 student 策略中;student 策略的奖励不含新颖性项。探索与优化交替进行,重复多轮。该解耦使得探索可以激进地扩大规模,而不损害 student 策略。在七个数学推理基准和两个模型族上,ExpDis 在相同 wall-clock 预算下优于 DAPO。此外,ExpDis 显示出更好的 pass@$k$ 扩展性,表明其生成的模型能够产出更多样化的正确解。
关键要点
- 01问题:在 RLVR 中加入强新颖性激励易损害模型质量,而可验证奖励仅监督狭窄的知识与行为,难以恢复。
- 02方法:提出 ExpDis 框架,将探索与优化解耦——探索者策略带新颖性奖励采样,过滤后蒸馏到无新颖性奖励的 student 策略,交替迭代多轮。
- 03结果:在 7 个数学推理基准、2 个模型族上,ExpDis 在相同 wall-clock 预算下优于 DAPO。
- 04结果:ExpDis 改善了 pass@$k$ 扩展性,生成更多样化的正确解,证明探索规模可激进放大而不退化 student。
解读
尚无解读。
原始英文摘要
arXiv:2610.10536v1 Announce Type: cross Abstract: Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@$k$ scaling, indicating that ExpDis produces models that generate more diverse correct solutions.