震耳欲聋的沉默:灾难性遗忘存在于数据从未提及的 Token 输出嵌入中
A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never Speaks
Jonghyun Han · Younghoon Song · Jongyoul Park
中文摘要
大语言模型(LLM)的持续预训练与微调不可避免地引发灾难性遗忘(catastrophic forgetting),通常依赖难以获取的原始数据进行回放(replay)缓解。在无数据条件下,对五种设置、最高 1.4B 参数模型进行系统性参数冻结分析发现:遗忘选择性集中于新语料中罕见 Token 的输出嵌入(output embedding);而主体中相同的 sqrt(v-hat) 区间保持惰性,新学习则位于其他区域。这种定位由语料的词表缺失(vocabulary deficiency)而非训练模式决定,因此可在固定基础模型内仅凭 Token 计数对预训练风险排序。缺失 Token 会持续受到单侧 softmax 梯度影响,Adam 二阶矩(sqrt(v-hat))归一化将其放大为完整更新。据此提出干预:训练期间仅提高输出投影层的 Adam epsilon。在八个设置、160M 至 12B 参数及四个模型家族中,该方法在全部七个稳定配置中消除 39.4% 至 67.9% 的遗忘,且不损害目标任务学习,无需逐模型调参。该防御与回放结合效果更佳(Qwen/Korean 达 79.8%),并能挽救已发布 head 的 LoRA,避免 23 倍遗忘激增。由于事后编辑漂移行只能恢复不足 5% 的遗忘,干预必须在训练期间实施。结果表明,当语料导致词表匮乏时,单行优化器调整即可作为抵御灾难性遗忘的主要防御手段。
关键要点
- 01无数据持续训练中的遗忘主要集中于罕见 Token 的输出嵌入,而非网络主体或由训练模式决定。
- 02缺失 Token 的单侧 softmax 梯度经 Adam 二阶矩归一化放大,形成完整参数更新。
- 03仅提高输出投影层的 Adam epsilon,在七个稳定配置中消除 39.4% 至 67.9% 的遗忘。
- 04干预无需逐模型调参,可与回放叠加使用,并在 Qwen/Korean 上达到 79.8%。
- 05事后编辑已漂移的输出行恢复不足 5%,说明防御必须作用于训练阶段。
解读
尚无解读。
原始英文摘要
arXiv:2610.09835v1 Announce Type: new Abstract: Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data. In this data-free regime, we analyze where forgetting occurs and why. Systematic parameter freezing across five settings up to 1.4B reveals that forgetting concentrates selectively in the output embeddings of tokens rarely seen in the new corpus, whereas the same sqrt(v-hat) band of the body is inert and new learning resides elsewhere. This localization is governed by the vocabulary deficiency of the corpus rather than the training mode, allowing pre-retraining risk ranking from token counts alone within a fixed base model. Mechanistically, absent tokens receive persistent one-sided softmax gradients that Adam's second-moment (sqrt(v-hat)) normalization amplifies into full-sized updates. We therefore propose an intervention: raising Adam's epsilon exclusively for the output projection during training. Across eight settings spanning 160M to 12B parameters and four model families, this removes 39.4% to 67.9% of forgetting across all seven stable configurations without degrading target learning or requiring per-model tuning. The defense combines additively or better with replay (79.8% on Qwen/Korean) and rescues released-head LoRA from a 23-fold forgetting surge. Because post-hoc editing of the drifted rows recovers under 5% of forgetting, the intervention must operate during training. Our findings indicate that a single-line optimizer adjustment may serve as the primary defense against catastrophic forgetting where the corpus starves the vocabulary.
同方向论文 · cs.CL
查看全部 →EngramEdit:通过条件记忆实现大语言模型的解耦知识更新
2610.10533Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响
2610.10508PHRBench:面向大语言模型幻觉后推理(后幻觉推理)的行为评测
2610.10455CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方
2610.10426使用大语言模型实现爱沙尼亚语文档级文本简化
2610.10378面向任务进度的行动学习:从紧凑教师监督中蒸馏小型智能体
2610.10332