通过直接最小化期望解码轮次训练并行投机解码草稿模型
Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds
Yunxiao Zhao · Changxiao Cai
中文摘要
投机解码利用低成本的草稿模型生成待验证的 token,由全尺寸目标模型并行验证,从而加速大语言模型推理。并行与半自回归(semi-AR)草稿模型能在单次前向中提出整块 token,提升草稿效率,但其训练带来新难题:某位置的草稿分布依赖于解码轮次从何处开始,而轮次的起点又取决于前几轮接受 token 的数量。现有训练目标多采用块内代理目标,忽略跨轮次耦合,因此无法直接优化全局解码效率。本文构建了将投机解码表示为马尔可夫奖励过程的理论与评估框架,推导出期望解码轮次(Expected Decoding Rounds, EDR)目标,该目标以状态占用率对局部拒绝代价加权,精确等于期望解码轮次数。EDR 不引入额外超参数。同时,基于时序差分推导出精确梯度,支持目标模型 rollout 的无偏随机优化。该框架还给出离线的精确轮次评估器,可在共享目标 rollout 上对比草稿模型,无需实际运行投机解码。在 DSpark 和 DFly 两种先进草稿模型上使用 EDR 微调,在九个数学推理、代码生成与对话基准上均稳定提升平均接受长度,优于现有训练目标。
关键要点
- 01问题:并行/半自回归草稿模型的训练目标为块内代理目标,忽略跨轮次耦合,无法直接优化全局解码效率
- 02方法:将投机解码建模为马尔可夫奖励过程,提出 EDR 目标精确等于期望解码轮次,并推导精确时序差分梯度支持无偏随机优化
- 03结果:在 DSpark 与 DFly 上用 EDR 微调,在九个数学推理、代码生成、对话基准上一致提升平均接受长度
- 04贡献:同一框架下给出精确离线轮次评估器,可在共享目标模型 rollout 上配对比较草稿模型,无需实际运行投机解码
- 05目标特性:EDR 不引入辅助超参数,直接以全局解码轮次为优化对象
解读
尚无解读。
原始英文摘要
arXiv:2610.10411v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel. Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted. Existing training objectives typically rely on block-local surrogates that ignore this cross-round coupling, and therefore do not directly optimize the global decoding efficiency. In this work, we develop a theoretical framework for training and evaluating these drafters by representing speculative decoding as a Markov reward process. This formulation yields the Expected Decoding Rounds (EDR) objective, which weights local rejection costs by state occupancies and exactly equals the expected number of decoding rounds. Unlike prior surrogate objectives, EDR introduces no auxiliary hyperparameters. We then derive an exact temporal-difference gradient that supports unbiased stochastic optimization from target-model rollouts. The same framework also yields an exact offline evaluator for round counts, enabling paired drafter comparisons on shared target rollouts without running speculative decoding. Finetuning two state-of-the- art drafters, DSpark and DFly, with EDR consistently improves mean accepted length and outperforms existing training objectives across nine benchmarks spanning math reasoning, code generation, and chat.