跳到正文
返回论文列表
cs.CL提交于 已译

EngramEdit:通过条件记忆实现大语言模型的解耦知识更新

EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

Hongru Cai · Ran Wei · Wenjie Wang · Chengfa Wu · Ning Song · Yongqi Li · et al.

中文摘要

条件记忆(conditional memory)架构(如 DeepSeek Engram)通过输入 n-gram 查找已学习的嵌入向量,以有限额外计算扩展大语言模型的参数容量。该架构不仅服务于模型扩展,还显示出了将事实性知识存储与通用计算解耦的可行性,为在不改主干 Transformer 的前提下更新事实知识提供了可能。然而实现这一目标存在难度:同一事实的不同表达方式会激活不同的 n-gram 嵌入,而更新共享嵌入可能意外改变模型对其他事实的预测。论文提出 EngramEdit,通过条件记忆实现解耦知识更新。EngramEdit 首先计算目标记忆表示,使模型在多种表达下均能预测更新后的事实;然后联合调整共享的 n-gram 嵌入,使其在多种表达和编辑下匹配这些表示,并对频繁复用的嵌入施加更强的惩罚以保留无关知识。实验表明 EngramEdit 可通过条件记忆实现独立的事实知识更新,编辑成功率接近完美。更新后的知识在未见过的表达以及多跳推理中均可用,在思维链(CoT)提示下的准确率约为最强基线的三倍。即便累积大量编辑,无关知识与通用能力也基本保留。这些发现表明 EngramEdit 将条件记忆转化为可编辑的知识接口,使其角色从模型扩展延伸至解耦知识更新。

关键要点

  1. 01问题:条件记忆中同一事实的不同表达激活不同 n-gram 嵌入,更新共享嵌入会意外改变其他事实的预测
  2. 02方法:先计算使多种表达均预测目标事实的记忆表示,再联合更新共享 n-gram 嵌入并对高频嵌入加强惩罚
  3. 03结果:编辑成功率接近完美;CoT 多跳推理准确率约为最强基线的三倍
  4. 04能力保留:累积事实编辑下,无关知识与模型通用能力基本不受影响

解读

尚无解读。

原始英文摘要

arXiv:2610.10533v1 Announce Type: new Abstract: Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory representations that make the model predict the updated fact across multiple expressions. It then jointly updates the shared n-gram embeddings to match these targets across expressions and edits, penalizing updates to frequently reused embeddings more strongly to preserve unrelated knowledge. Experiments show that EngramEdit enables independent factual knowledge updates through conditional memory, achieving near-perfect editing success. Revised knowledge is usable across unseen expressions and in multi-hop reasoning, with nearly three times the strongest baseline's accuracy under chain-of-thought (CoT) prompting. Unrelated knowledge and general capabilities are largely preserved even as factual updates accumulate. These findings show that EngramEdit turns conditional memory into an editable knowledge interface, extending its role beyond model scaling to support decoupled knowledge updates.

同方向论文 · cs.CL

查看全部 →