面向交错多模态生成的自校正优化
Self-correction Optimization for Interleaved Multimodal Generation
Xin You · Zhiwei Ning · Zukai Chen · Minghui Zhang · Xuanke Shi · Hanxiao Zhang · et al.
中文摘要
多模态大语言模型(MLLMs)在视觉理解与生成方面取得显著进展,但生成图文交错内容仍是难题,因为其需要紧密耦合的多模态理解与生成能力。现有 MLLM 虽提供了可行方案,大多通过增广数据进行额外训练,计算开销大,且在视觉主体保持、时间一致性以及物理合理性方面仍存在局限。本文提出自校正优化(SCO),一种无需训练即可实现一致性交错生成的有效方法。SCO 将无分类器引导(classifier-free guidance)的更新作为参考,在新事件约束与状态保持约束两类互补约束下进行最小化自校正:新事件约束促进图文序列间的时间一致性,状态保持约束在后续生成步骤中维持视觉主体的连贯性。在具有挑战性的交错多模态生成 benchmark 上的实验表明,SCO 在时间连贯性与视觉主体保持上带来显著提升。此外,SCO 可扩展至视频生成,改善对物理基础过程的建模,包括机器人操控与长程手工艺任务。
关键要点
- 01问题:交错图文生成需要紧密耦合的理解与生成能力,现有 MLLM 依赖增广数据额外训练,计算昂贵,在主体保持、时间一致性与物理合理性上仍有不足
- 02方法:SCO 将 classifier-free guidance 的更新视作参考,在新事件约束与状态保持约束下做最小化自校正,无需额外训练
- 03结果:在交错多模态生成 benchmark 上显著提升时间连贯性与视觉主体保持
- 04扩展:可迁移至视频生成,改善机器人操控与长程手工艺等物理基础过程的建模
- 05优势:训练-free,避免增广数据与额外训练的计算开销
解读
尚无解读。
原始英文摘要
arXiv:2610.10400v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made significant progress in visual understanding and generation. However, generating interleaved image--text content remains challenging, as it requires tightly integrated multimodal understanding and generation capabilities. Although existing MLLMs provide promising solutions, most rely on additional training with augmented data, which is computationally expensive and remains limited in preserving visual subjects, temporal consistency, and physical plausibility. In this work, we propose self-correction optimization (SCO), an effective training-free method for consistent interleaved generation. SCO treats the classifier-free guidance update as a reference and performs minimal self-correction under two complementary constraints, including new-event and state-preserving constraints. Specifically, the new-event constraint promotes temporal consistency across image--text sequences, while the state-preserving constraint maintains the coherence of visual subjects throughout subsequent generation steps. Experiments on challenging interleaved multimodal generation benchmarks demonstrate significant improvements in temporal coherence and visual-subject preservation. Furthermore, SCO can be extended to video generation and improves the modeling of physically grounded processes, including robot manipulation and long-horizon handcrafting.