跳到正文
返回论文列表
cs.LG提交于 已译

ResidualQuant:面向循环 Transformer 的 2 比特残差键值缓存(KV cache)量化方法

ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals

Heejun Kim · Junyoung Lee · SangLyul Cho · Dongsu Han · Insu Han · Sehoon Kim

中文摘要

循环 Transformer(Looped Transformers)通过在多个循环迭代中重复应用共享 Transformer 模块来增加计算深度,从而改善参数利用率。但键值缓存(Key-Value cache,简称 KV cache)内存随循环数线性增长,成为限制批大小和推理吞吐的关键内存瓶颈。键值缓存量化可以缓解该瓶颈,但现有方法在激进低精度下往往出现显著的精度下降。我们观察到循环 Transformer 提供了一个独特机会:各循环间的键值状态高度相似。基于此提出 ResidualQuant,以最后一轮的键值状态为参考,将其余各轮表示为低精度残差。该方法进一步结合最小二乘缩放与旋转对残差进行处理,并采用按循环混合精度策略,在保留高效重建的前提下实现 INT2 精度的准确量化。在多种循环 Transformer 模型及数学推理与代码生成基准上,ResidualQuant 始终在精度-内存权衡上优于当前最优的基于旋转的键值缓存量化方法。具体而言,在混合精度设置下,该方法在理论键值存储降低 80.7% 的同时保留接近 BF16 的精度,在相同内存预算下比基于旋转的基线最高提升 13.0% 的精度。在 RTX 5090 上,减少的键值内存流量使固定批大小的解码吞吐最高提升 2.73 倍;同时更小的内存占用支持最高 2 倍的批大小,峰值吞吐最高提升 4.15 倍。

关键要点

  1. 01问题:循环 Transformer 的键值缓存随循环数线性增长,成为限制批大小和推理吞吐的内存瓶颈,现有低精度量化方法精度损失严重
  2. 02方法:利用各循环键值状态的高度相似性,以最后一轮键值状态为参考,将其余轮编码为低精度残差,并结合最小二乘缩放、旋转及按循环混合精度,将量化精度下探至 INT2
  3. 03结果:在多个循环 Transformer 模型与数学推理、代码生成基准上,精度-内存权衡优于基于旋转的键值缓存量化基线,混合精度下理论键值存储降低 80.7% 并保留接近 BF16 的精度,相同内存预算下精度最高提升 13.0%
  4. 04吞吐收益:在 RTX 5090 上,固定批大小解码吞吐最高提升 2.73 倍,峰值吞吐因批大小最高翻倍而提升 4.15 倍
  5. 05局限:摘要未讨论在更长循环数或非数学/代码任务上的退化情况及与其它非量化压缩方法的比较

解读

尚无解读。

原始英文摘要

arXiv:2610.10381v1 Announce Type: new Abstract: Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size and inference throughput. KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes. We observe that looped Transformers offer a unique opportunity: KV states across loops are highly similar. Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops with low-precision residuals. Our method further combines least-square scaling and rotations applied to the residuals, as well as loop-wise mixed precision, to enable accurate quantization down to INT2 while retaining efficient reconstruction. Across multiple looped Transformer models and mathematical reasoning and code generation benchmarks, ResidualQuant consistently improves the accuracy-memory tradeoff over state-of-the-art rotation-based KV quantization. In particular, our method retains accuracy close to BF16 under mixed-precision settings while reducing theoretical KV storage by 80.7%, achieving up to 13.0% higher accuracy than the rotation-based baseline at the same memory budget. On an RTX 5090, the reduced KV memory traffic improves fixed-batch decode throughput by up to 2.73x, while the smaller memory footprint enables up to 2x larger batches, improving peak throughput by up to 4.15x.

同方向论文 · cs.LG

查看全部 →