跳到正文
返回论文列表
cs.LG提交于 已译

基于 Koopman 观测器的扩散加速:用浅层测量修正特征预测

Koopman Observers for Diffusion Acceleration: Correcting Feature Forecasts with Shallow Measurements

Hanru Bai · Yuanchao Xu · Fengyi Li

中文摘要

特征缓存(feature caching)通过用先前计算的激活预测来替代昂贵的网络评估,从而加速扩散采样。然而,仅基于过去特征的预测无法直接融合当前去噪状态的变化。本文研究廉价的、即时计算的浅层特征能否作为观测信号以修正这些预测。提出一种观测修正的 Koopman 框架,用于加速冻结的扩散模型。通过标定轨迹,识别出有限维、时变的 Koopman 近似,联合描述浅层与深层网络特征的增量。在加速采样阶段,这些算子预测昂贵深层特征的演化,同时观测到的浅层特征的新息(innovation)对预测状态进行修正。周期性的完整网络评估用于刷新观测器,所有生成模型参数保持不变。该框架支持对时间预测与观测修正进行受控比较。在每个数据集的三轮各 10,000 张图像实验中,相对同一四步部分时间步进下的逐通道仿射预测方法,本文方法将配对的 Inception 特征 MSE 分别降低 19.9%(CIFAR-10)和 11.9%(ImageNet 十类子集)。匹配消融实验显示,观测修正分别带来额外 4.54% 与 4.67% 的降幅。观测器相对 DDIM-50 取得 1.89 倍与 1.85 倍的实测加速,在无需重训去噪器的条件下提升参考采样器的保真度。

关键要点

  1. 01问题:现有特征缓存只用历史激活预测深层特征,忽略当前去噪状态的即时变化
  2. 02方法:构建时变有限维 Koopman 算子联合建模浅层与深层特征增量,并以浅层新息在线修正深层预测
  3. 03结果:CIFAR-10 与 ImageNet 十类子集上配对 Inception 特征 MSE 分别降低 19.9% 和 11.9%,相对 DDIM-50 取得约 1.85–1.89 倍实测加速
  4. 04局限:观测器需周期性完整评估刷新,且只在四步部分时间步进 schedule 下与通道仿射基线做了受控对比

解读

尚无解读。

原始英文摘要

arXiv:2610.10366v1 Announce Type: new Abstract: Feature caching accelerates diffusion sampling by replacing expensive network evaluations with predictions from previously computed activations. However, forecasts based only on past features cannot directly incorporate changes in the current denoising state. We investigate whether inexpensive, freshly computed features can serve as observations for correcting these predictions. We introduce an observation-corrected Koopman framework for accelerating frozen diffusion models. Using calibration trajectories, we identify finite-dimensional, time-dependent Koopman approximations that jointly describe the increments of shallow and deep network features. During accelerated sampling, these operators predict the evolution of expensive deep features, while innovations in the observed shallow features correct the predicted state. Periodic full evaluations refresh the observer, and all generative-model parameters remain unchanged. This formulation enables controlled comparisons of temporal prediction and observation correction. Across three 10,000-image runs per dataset, our method reduces paired Inception-feature MSE by $19.9\%$ on CIFAR-10 and $11.9\%$ on a ten-class ImageNet subset relative to channelwise affine prediction under the same four-partial-step schedule. Matched ablations attribute additional reductions of $4.54\%$ and $4.67\%$ to observation correction. The observer achieves $1.89\times$ and $1.85\times$ measured speedups over DDIM-50, supporting improved reference-sampler fidelity without retraining the denoiser.

同方向论文 · cs.LG

查看全部 →