跳到正文
返回论文列表
cs.CV提交于 已译

MORCA:面向视频扩散加速中自适应缓存复用的离线到在线强化学习

MORCA: Offline-to-Online Reinforcement Learning for Adaptive Cache Reuse in Video Diffusion Acceleration

Yuxiang Xiong · Ruiyan Wang · Wenqiang Wang · Teng Hu · Songhang Shen · Bohao Feng · et al.

中文摘要

扩散 Transformer(DiT)在视频合成中表现优异,但其迭代去噪过程导致推理延迟较高。为解决此问题,缓存作为一种有效的加速策略被提出,利用去噪过程中的跨步冗余。现有动态缓存方法通常估计每一步缓存复用所引入的误差(逐步误差),而本文关注的是缓存复用对最终生成视频造成的质量损失(终端误差)。研究表明,逐步误差与终端误差并不直接对应,潜在信息有助于捕捉二者关系,从而指导缓存决策。此外,现有基于阈值的方法无法提供精确的加速控制,难以满足用户指定的实际加速目标。针对上述局限,本文提出 MORCA,一种通过离线到在线强化学习训练的缓存调度框架,在用户指定的加速目标下进行潜在感知的复用与重计算决策。在不同视频生成模型与多种目标加速比上的大量实验表明,在相近的计算预算下,MORCA 比现有最优缓存方法实现了更高的生成保真度。代码开源于 https://github.com/x10ngyx/MORCA。

关键要点

  1. 01问题:视频 DiT 迭代去噪延迟高,现有动态缓存依据逐步误差做判断且阈值法无法精确控制加速比,难以满足用户指定加速目标
  2. 02方法:提出 MORCA 框架,基于离线到在线强化学习,在用户指定加速目标下利用潜在表示(latents)进行复用与重计算的调度决策
  3. 03结果:在不同视频生成模型与多档目标加速比下,MORCA 在相近算力预算内取得优于现有最优缓存方法的生成保真度
  4. 04价值:把缓存决策从逐步误差转向终端误差,并支持精确的速度调节,更贴合实际部署中的加速需求

解读

尚无解读。

原始英文摘要

arXiv:2610.10457v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) achieve remarkable performance in video synthesis, but their iterative denoising process suffers from high inference latency. To address this, caching has emerged as an effective acceleration strategy by capitalizing on inter-step redundancy during denoising. Existing dynamic caching methods typically estimate the error that cache reuse would introduce at each denoising step (step error) to guide cache decisions, whereas our concern is how much quality loss cache reuse would cause in the final generated video (terminal error). We show that step error does not directly correspond to terminal error and that latent information helps capture their relationship, thereby informing cache decisions. Moreover, existing threshold-based methods cannot provide precise speedup control, making it difficult to meet practical requirements for user-specified acceleration targets. To address these limitations, we introduce MORCA, a cache scheduling framework trained through offline-to-online reinforcement learning to make latent-aware reuse/recompute decisions under user-specified acceleration targets. Extensive experiments on different video generation models across multiple target acceleration ratios demonstrate that MORCA achieves better generation fidelity than state-of-the-art caching methods under comparable computational budgets. Code is available at https://github.com/x10ngyx/MORCA.

同方向论文 · cs.CV

查看全部 →