跳到正文
返回论文列表
cs.CL提交于 已译

在潜空间中评判:通过语义保持压缩实现高效的生成式奖励建模

Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression

Mingqing Yuan (Soochow University) · Xiaobo Liang (Soochow University) · Junwei Yang (University of Cambridge) · Ziwei Chen (Chalmers University of Technology) · Zeren Zhang (Peking University) · Hejin Wang (Tsinghua University) · et al.

中文摘要

奖励建模通常需要对多个评估准则进行联合表示与推理,但逐 token 的文本化过程会带来显著的推理开销。近期关于潜空间推理(latent reasoning)的研究表明,连续状态可以更紧凑地承载此类计算。本文提出 LatentGRM,一种基于语义分块、压缩与重建的潜空间评估框架。LatentGRM 利用评分量表(rubric)引导的评估结构来指导压缩,学习到紧凑的连续轨迹,从而在无需生成文本评估的前提下自主完成成对比较判断。一个独立的解释器(interpreter)从这些轨迹中重建评估文本,提供压缩下信息保留情况的离线视角。在训练数据与骨干网络一致的条件下,LatentGRM 在 4B 与 8B 规模下均取得了与显式监督微调(SFT)评判器相当的聚合偏好准确率。在四个基准领域上,LatentGRM-8B 将评估轨迹压缩 8.9–9.2 倍,并在 vote@5 下将评判器总推理时间降低 6.1–7.0 倍。受控的评分量表干预实验表明,依赖具体准则的偏好信息能够通过潜序列传递。上述结果表明,连续潜空间评估在显著降低推理成本的同时,能够保持具有竞争力的判断质量。

关键要点

  1. 01问题:奖励建模中逐 token 的文本化推理开销大,连续潜空间有望更紧凑地承载多准则评估计算
  2. 02方法:LatentGRM 利用评分量表结构引导语义分块、压缩与重建,学习紧凑的连续评估轨迹进行成对判断,并配独立解释器离线重建文本
  3. 03结果:在四个基准上,LatentGRM-8B 压缩评估轨迹 8.9–9.2 倍,vote@5 下推理时间降低 6.1–7.0 倍,偏好准确率与显式 SFT 评判器相当
  4. 04验证:受控的评分量表干预表明,依赖具体准则的偏好信息在潜序列中得以保留
  5. 05规模:在 4B 与 8B 规模下均有效,匹配训练数据与骨干网络

解读

尚无解读。

原始英文摘要

arXiv:2610.09788v1 Announce Type: new Abstract: Reward modeling often requires jointly representing and reasoning over multiple evaluation criteria, yet verbalizing this process token by token can incur substantial inference cost. Recent work on latent reasoning suggests that continuous states may support this computation more compactly. We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction. By using the structure of rubric-guided evaluations to guide compression, LatentGRM learns compact continuous trajectories that support autonomous pairwise judgments without generating textual assessments. A separate interpreter reconstructs evaluation text from these trajectories, providing an offline view of the information retained under compression. Under matched training data and backbones, LatentGRM achieves competitive aggregate preference accuracy relative to explicit Supervised Fine-Tuning (SFT) judges at both 4B and 8B scales. Across four benchmark domains, LatentGRM-8B compresses evaluation trajectories by 8.9--9.2x and reduces total judge inference time by 6.1--7.0x at vote@5. Controlled rubric interventions show that criterion-dependent preference information is carried through the latent sequence. Together, these results demonstrate that continuous latent evaluation can substantially reduce inference cost while preserving competitive judgment quality.

同方向论文 · cs.CL

查看全部 →