跳到正文
返回论文列表
cs.AI提交于 已译

Few Bits, One Law:面向 W2A4KV2 的统一量化

Few Bits, One Law: Toward W2A4KV2

Kai Yi · Tarek Elgamal · Sruthikesh Surineni · Vignesh Vivekraja · Soumyadeep Ghosh · Steven Li

中文摘要

极低比特 LLM 压缩最具挑战性的情形是权重、激活值与 KV 缓存同时量化:三者的分布各异,且量化误差在网络中相互耦合。CanonQ 是一个统一的量化感知训练框架,通过将源规范化(lossy normalization)与任务感知自适应解耦来应对这些挑战。固定的旋转矩阵与能量归一化将异构张量源映射到规范坐标,使冻结的高斯参考码本可在层间和模型间复用。联合训练随后在统一的标量/向量接口下,使网络适配权重、激活值与缓存量化耦合产生的误差。论文给出了冻结码本迁移误差与局部任务损失的界,并推导出精确的归一化感知直通估计 Jacobian,将量化失真与梯度偏差联系起来。最显著的收益出现在联合 W2A4KV2 压缩下:在 LLaMA3-1B/3B/8B 上,CanonQ-Omni 相对此前 SOTA 及代表性量化基线,WikiText-2 困惑度最多降低 14.28 倍,零样本平均准确率最多提升 57.9%。该优势还扩展到 Qwen3-1.7B、代码生成与数学推理:在指令微调的 MobileLLM-Pro-1B 上 W2A16KV16 配置下,CanonQ 相对最强量化基线,HumanEval pass@1 相对提升 41.7%,GSM8K exact match 相对提升 39.1%。

关键要点

  1. 01问题:权重、激活值与 KV 缓存同时量化时,分布差异与误差耦合使极低比特 LLM 压缩极具挑战
  2. 02方法:CanonQ 通过固定旋转和能量归一化将异构张量映射到规范坐标,复用冻结高斯码本,并在统一标量/向量接口下做任务感知联合训练
  3. 03结果:在 LLaMA3-1B/3B/8B 的 W2A4KV2 配置下,WikiText-2 困惑度最多降低 14.28 倍,零样本准确率最多提升 57.9%,并在 Qwen3-1.7B 与 MobileLLM-Pro-1B 上同样取得 SOTA
  4. 04理论:给出了冻结码本迁移误差与局部任务损失的界,并推导出归一化感知直通估计的精确 Jacobian,将量化失真与梯度偏差联系起来

解读

尚无解读。

原始英文摘要

arXiv:2610.09202v1 Announce Type: cross Abstract: Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation. Fixed rotations and energy normalization map heterogeneous tensor sources to canonical coordinates, enabling frozen Gaussian-reference codebooks to be reused across layers and models. Joint training then adapts the network to the coupled errors of weight, activation, and cache quantization within a common scalar/vector interface. We bound frozen-codebook transfer error and local task loss, and derive an exact normalization-aware straight-through Jacobian that links quantization distortion to gradient bias. The strongest gains arise under joint W2A4KV2 compression: across LLaMA3-1B/3B/8B, CanonQ-Omni achieves up to 14.28x lower WikiText-2 perplexity and up to 57.9% higher mean zero-shot accuracy than prior state-of-the-art and representative quantization baselines. The benefits extend to Qwen3-1.7B, code generation, and mathematical reasoning: on instruction-tuned MobileLLM-Pro-1B at W2A16KV16, CanonQ achieves relative improvements of 41.7% in HumanEval pass@1 and 39.1% in GSM8K exact match over the strongest evaluated quantization baseline.

同方向论文 · cs.AI

查看全部 →