OrBIT:基于结构引导的嵌入压缩
OrBIT: Structure-Guided Embedding Compression
Yunied Puig · Amit Kumar Jaiswal
中文摘要
嵌入表(embedding table)是现代语言模型中体积最大的组件之一。多数压缩方法预先固定编码几何,例如坐标分块、低秩子空间或无约束码本,然后在该框架内优化。本文转而探讨编码几何本身能否被发现。提出 OrBIT,一种结构引导的嵌入压缩框架,通过轨道动力学(orbit dynamics)学习可复用的局部几何,并以此约束一组少量共享码字。全局重构残差决定固定编码预算的分配位置,冗余的局部重叠图表在拼接后使局部误差相互补偿。理论分析表明:紧致图表几何如何控制失真,全局残差如何指导顺序分配,以及数据几何引导的细化如何改进编解码器。学到的轨道机制在编译时被消除,最终得到一个紧凑的解码器,由学到的结构决定存储内容、容量分配与局部信息的全局组装方式。在四个 LLM 嵌入表上,OrBIT 在 GPT-2 上达到 37.9× 压缩率,在每个 7B 表上相对 16-bit 存储达到 23× 以上压缩率,并在率-失真曲线上与主流量化和低秩基线方法保持竞争力。
关键要点
- 01问题:现有嵌入压缩方法固定编码几何后在框架内优化,编码几何本身是否可被自动发现仍是开放问题。
- 02方法:OrBIT 通过轨道动力学学习可复用局部几何,约束共享码字集合,再由全局残差指导预算顺序分配。
- 03结果:在 GPT-2 上达到 37.9× 压缩,各 7B 模型嵌入表上超过 23× 压缩,率-失真性能优于或持平量化与低秩基线。
- 04理论:给出紧致图表几何的失真控制、全局残差的顺序分配指导,以及数据几何引导的细化对编解码器的改进。
- 05局限:轨道机制在编译时被消除,最终仅保留紧凑解码器,具体推理开销与部署兼容性未在摘要中说明。
解读
尚无解读。
原始英文摘要
arXiv:2610.10385v1 Announce Type: new Abstract: Embedding tables are among the largest components of modern language models. Most compression methods fix a coding geometry such as coordinate blocks, low-rank subspaces, or unrestricted codebooks, and optimize within it. We instead ask whether the coding geometry can itself be discovered. We introduce \emph{OrBIT}, a structure-guided embedding compression framework that learns reusable local geometry from orbit dynamics and uses it to constrain a small set of shared codewords. The global reconstruction residual then decides where the fixed coding budget is spent, while redundant overlapping charts let local errors compensate one another after gluing. Our theory shows how tight-chart geometry controls distortion, how the global residual directs sequential allocation, and how data-geometry-guided refinement improves the codec. The resulting orbit machinery is compiled away, leaving a compact decoder in which the learned structure governs what is stored, where capacity is allocated, and how local information is assembled globally. Across four LLM embedding tables, OrBIT achieves $37.9\times$ compression on GPT-2 and over $23\times$ on each 7B table relative to 16-bit storage, while delivering competitive rate-distortion performance against established quantization and low-rank baselines.