长上下文混合模型的机制 第1.1部分:从混合注意力到混合位置
Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position
Xiaoran Liu · Ziwei He · Xipeng Qiu
中文摘要
大语言模型(LLMs)的架构设计正从传统的纯全注意力模型转向混合模型,后者组合多种注意力模块以提升长上下文效率以及长度外推和上下文扩展的性能。为解释混合模型为何有效以及如何更好地设计,本研究提出《长上下文混合模型的机制》。作为系列第1.1部分,从全注意力与滑动窗口注意力(SWA)或门控线性注意力(LA)变体(包括 GLA 和 GDN)的混合模型入手。研究首先观察到上下文扩展中的跷跷板效应:LA 混合模型从长上下文持续预训练中获益更多,而 SWA 混合模型在长度外推下表现更优。该行为归因于不同注意力机制所引入的位置归纳偏置差异。研究发现 SWA 混合模型存在短上下文学习陷阱、短窗口倦怠与长窗口惰性问题,需要扩展窗口以提升持续长上下文预训练中的性能。对于 LA 混合模型,归纳出混合位置外推的马太效应,并提出滑动窗口线性注意力,在 64k 上下文长度下实现 16 倍免训练长度外推,同时在 NIAH-SK1 上保持 100% 准确率。
关键要点
- 01问题:全注意力模型在长上下文效率与长度外推上存在局限,需要从架构层面转向混合注意力设计
- 02方法:系统性比较全注意力与 SWA、GLA、GDN 等混合变体,分析位置归纳偏置差异,并提出滑动窗口线性注意力
- 03发现:上下文扩展中存在跷跷板效应——LA 混合模型在持续预训练中更强,SWA 混合模型在长度外推中更优
- 04局限:SWA 混合模型面临短上下文学习陷阱、短窗口倦怠与长窗口惰性,需扩展窗口才能改善性能
- 05结果:所提滑动窗口线性注意力在 64k 上下文下实现 16 倍免训练长度外推,NIAH-SK1 准确率 100%
解读
尚无解读。
原始英文摘要
arXiv:2610.10114v1 Announce Type: new Abstract: The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation. We attribute this behavior to differences in the positional inductive biases induced by these attention mechanisms. We find that SWA hybrids suffer from a Short-Context Learning Trap, Short-Window Weariness, and Long-Window Laziness, and require extended windows to enhance performance in continual long-context pretraining. For LA hybrids, we summarize the Matthew Effect of Hybrid Position Extrapolation and propose Sliding-Window Linear Attention, achieving 16$\times$ training-free length extrapolation while maintaining 100\% accuracy on NIAH-SK1 in 64k context length.
同方向论文 · cs.CL
查看全部 →EngramEdit:通过条件记忆实现大语言模型的解耦知识更新
2610.10533Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响
2610.10508PHRBench:面向大语言模型幻觉后推理(后幻觉推理)的行为评测
2610.10455CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方
2610.10426使用大语言模型实现爱沙尼亚语文档级文本简化
2610.10378面向任务进度的行动学习:从紧凑教师监督中蒸馏小型智能体
2610.10332