SeOPD:通过自生成思维链在线策略蒸馏实现自我进化 LLM
SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought
Xiaoshu Chen · Sihang Zhou · Ke Liang · Xinwang Liu
中文摘要
在线策略自蒸馏(OPSD)的最新研究表明,大语言模型(LLM)可利用外部特权信息(PI)(如人工标注或外部环境反馈)提升能力,但获取准确标注和构建复杂环境需要大量人力与算力,限制其扩展性。近期虽有无外部 PI 的自我改进研究,增益仍有限。单个 LLM 可支持深度思考与非思考等多种推理模式,深度思考能在推理中生成额外信息。自我进化在线策略蒸馏(SeOPD)让 LLM 蒸馏并内化自身生成的思维链(CoT):先用深度思考模式生成 CoT,再用非思考模式生成回答,并将 CoT 作为 PI,为非思考回答提供 token 级监督,使推理得到的新信息指导非思考模式,并写入共享模型参数,从而同时提升两种模式。跨多个 LLM 与任务的实验验证了其有效性。
关键要点
- 01OPSD依赖人工标注或外部反馈,获取成本高且扩展受限
- 02核心观察:单个LLM可运行深度思考与非思考等推理模式
- 03深度思考CoT可作为特权信息,为非思考回答提供token级监督
- 04推理产生的新信息被内化至共享参数,同时提升两种模式能力
- 05跨LLM与任务的广泛实验验证了SeOPD的有效性
解读
尚无解读。
原始英文摘要
arXiv:2609.33181v2 Announce Type: replace-cross Abstract: Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.