SGF+:解耦自回归视频生成的梯度流
SGF+: Decoupling Gradient Flows for Autoregressive Video Generation
Zihan Su · Junhao Zhuang · Yaowei Li · Siwen Lu · Haoran Li · Lingen Li · et al.
中文摘要
自回归视频生成在预测未来帧时,既要对当前帧去噪,也需写入键值表示作为上下文。然而,这两种角色通常共享参数,其梯度呈现不同模式并存在系统性的负对齐,妨碍视觉质量与时序一致性的联合优化。Self Gradient Forcing Plus(SGF+)为上下文写入与去噪分配独立参数,同时通过因果注意力保持二者交互;两种角色均使用原始生成目标联合优化,不引入辅助损失,并以上下文写入对未来预测的贡献提供监督。该方法在逐帧与分块生成中均提升了视觉质量和长程一致性,无需额外视频训练数据或延长训练周期。仅在 5s 滚动序列上训练后,SGF+ 无需长视频微调即可连续生成最长 24 小时的视频。结果表明,面向角色的参数化是实现高质量自回归视频生成与原生长程外推的有效设计原则。
关键要点
- 01共享参数导致上下文写入与去噪的梯度系统性负对齐,妨碍联合优化。
- 02SGF+为上下文写入和去噪分配独立参数,并通过因果注意力保持交互。
- 03方法仅使用原始生成目标联合优化,不依赖辅助损失。
- 04在逐帧与分块生成中,视觉质量和长程一致性均优于所评估基线。
- 05仅训练5s滚动序列,即可连续生成最长24小时视频,无需长视频微调。
解读
尚无解读。
原始英文摘要
arXiv:2610.10429v1 Announce Type: new Abstract: Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint optimization of visual quality and temporal consistency. We introduce Self Gradient Forcing Plus (SGF+), which assigns separate parameters to context writing and denoising while preserving their interaction through causal attention. Both roles are jointly optimized using the original generation objective without auxiliary losses, with context writing supervised through its contribution to future predictions. This simple change improves visual quality and long-horizon consistency over the evaluated baselines in both framewise and chunkwise generation, without additional video training data or a longer training horizon. Trained on only 5s rollouts, SGF+ supports continuous generation for up to 24 hours without long-video fine-tuning. These results highlight role-specific parameterization as an effective design principle for high-quality autoregressive video generation and native long-horizon extrapolation.