跳到正文
返回论文列表
cs.CL提交于 已译

缓存编码器:跨 LLM 查询的紧凑可复用记忆

Cache the Encoder Within:Compact, Reusable Memory across LLM Queries

Hanzuo Liu · Chunyu Liu · Chaofan Lin · Alex Lamb · Mingyu Gao

中文摘要

arXiv:2610.10058v1 公告类型:新论文。针对共享文档的重复查询会产生冗余编码,而缓存模型状态又会带来持续存储成本。EncBank 基于 CoMem 的中间状态接口,将预训练 LLM 的较低层用作可复用文档编码器,紧凑存储其输出,供适配后的上层读取器使用。自蒸馏后缀适配器在同一骨干网络的不同存储精度间共享,无需针对量化重新训练。在三个 Qwen 骨干网络、五个 benchmark 套件上,涵盖不同规模以及全注意力和混合架构;4-bit 存储使各 benchmark 汇总分数与原生精度 EncBank 的差距不超过 1 分。在固定的 Qwen3-8B 工作负载中,其持久 GPU 存储占原生精度的 28.1%。原生精度对照实验相较相同证据、相同适配器的文本重放,实现选定 pack 的 prefill 1.40x 加速,但 RULER 准确率下降 3.12 分。原生精度 Qwen3.8-27B 配置通过 Terminal-Bench 2.1 的 89 项任务中的 70 项。EncBank 因而结合了可复用计算与紧凑记忆,但任务保真度和端到端收益仍取决于工作负载、准备成本与复用频率。

关键要点

  1. 01针对共享文档重复查询的冗余编码与状态缓存存储成本,EncBank 将 LLM 低层输出存为可复用文档编码结果。
  2. 02自蒸馏后缀适配器跨存储精度共享,避免量化专用重训,并支持适配后的上层读取器。
  3. 03在三个 Qwen 骨干、五个 benchmark 套件上,4-bit 存储与原生精度的汇总分数差距不超过 1 分。
  4. 04Qwen3-8B 固定工作负载仅保留原生精度持久 GPU 存储的 28.1%,原生精度文本重放对照的选定 pack prefill 加速 1.40x。
  5. 05RULER 准确率下降 3.12 分;Qwen3.8-27B 通过 Terminal-Bench 2.1 的 70/89 项任务。

解读

尚无解读。

原始英文摘要

arXiv:2610.10058v1 Announce Type: new Abstract: Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.

同方向论文 · cs.CL

查看全部 →