GRACE:面向生成的高效视频生成潜空间压缩方法
GRACE: Generation-aware latent compression for efficient video generation
Jiyoung Kim · Paul Hyunbin Cho · Jisu Nam · Donghoon Lee · Hyunsung Go · Yeonkyeong Lee · et al.
中文摘要
高压缩率视频自编码器可通过大幅减少 token 数量加速视频扩散模型。然而,压缩比升高会降低重建质量,恢复质量需要更多通道,而更多通道会拖慢 DiT(Diffusion Transformer)收敛,且训练成本高昂,自编码器的潜空间也偏离 DiT 原始训练分布。对 DiT 原训练所用的自编码器做压缩虽能保持兼容性,但仅以重建为目标的优化仍会改变潜空间分布。为此,论文提出 GRACE(Generation-Aware Latent Compression for Efficient Video Generation),一个两阶段框架,可在保持与预训练 DiT 兼容的前提下压缩预训练视频自编码器。方法冻结基潜空间,学习残差潜空间以补充强压缩下丢失的信息,并在冻结 DiT 特征空间中对齐压缩后的潜空间与原始潜空间,使自编码器针对生成优化。随后对 DiT 做轻量微调并采用非对称去噪流程,先去噪基再处理残差。GRACE 将 Wan2.1-I2V-14B 的 token 数量降至 1/8,在 480×832×81 分辨率下延迟降低 11.1×,在 VBench 上生成质量与压缩前预训练管线持平。
关键要点
- 01问题:对视频自编码器做强压缩会降低重建质量、增加通道数、拖慢 DiT 收敛,并使潜空间偏离 DiT 训练分布,重新训练成本高。
- 02方法:GRACE 采用两阶段框架,冻结原编码器的基潜空间,学习残差潜空间补足丢失信息,并在冻结 DiT 特征空间中对齐压缩潜空间与原潜空间。
- 03推理策略:在去噪时采用非对称流程,先对基去噪再处理残差,并对 DiT 进行轻量微调以适配压缩后的潜空间。
- 04结果:在 Wan2.1-I2V-14B 上 token 数降至 1/8、延迟降低 11.1×(480×832×81),VBench 生成质量与压缩前持平。
- 05局限:论文摘要未涉及明显局限;该方法依赖预训练 DiT 与自编码器对,且与未公开的内部基准未做对比。
解读
尚无解读。
原始英文摘要
arXiv:2610.10524v1 Announce Type: new Abstract: Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT. Specifically, we keep a frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression, while aligning the compressed latent with the pretrained latent in the feature space of the frozen DiT so that the autoencoder is optimized for generation. We then adapt the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual. GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.