并行适配器组合实现实时联合音视频生成
Real-Time Joint Audio-Video Generation by Parallel Adapter Composition
Jingyu Li · Xiaoxiao Xiang · Yiwen Guo
中文摘要
面向实时交互式生成的联合音视频扩散变换器通常需要两项关键改造:块自回归注意力,使帧在整段视频完成前即可输出;少步采样,使每个块计算开销低。传统流式视频方案通过链式管线获得这两项能力——先蒸馏双向教师为因果学生,再蒸馏为少步模型,或顺序相反。该链路的每一阶段都对前阶段产出的权重进行微调,后期目标可能破坏既有能力。受模型融合思路启发,本文证明在打包的音视频主干上,两项能力可并行获取:在冻结的主干上训练一个因果适配器,并复用现成的少步适配器提供少步能力。由于二者修改的是不同的功能轴,论文预测并验证其权重更新方向近似正交,训练中未施加任何显式正交约束。正交更新可直接叠加而不互相干扰,因此并行组合即为直接求和。推理时两个适配器简单相加,无需联合训练,即可得到与双向教师图像质量相当的少步流式音视频生成。相比链式基线,组合模型在多数指标上持平或更优,证明并行组合的实用性。最终流式系统可实时生成联合音视频,480×832 分辨率下约 26 fps(无需量化),并能在 30 秒内持续生成且图像质量稳定。
关键要点
- 01问题:传统链式蒸馏方法在将双向扩散模型改造为流式、少步模型时,后一阶段的微调可能破坏前阶段已获得的能力。
- 02方法:在冻结的打包音视频主干上并行训练因果适配器与现成少步适配器,利用模型融合思想实现直接相加组合。
- 03原理:因果适配器与少步适配器作用于不同功能轴,其权重更新方向近似正交,训练中无需显式正交约束即可无干扰叠加。
- 04结果:组合模型在多数指标上匹配或优于链式基线,图像质量与双向教师相当,且无需联合训练。
- 05性能:系统在 480×832 分辨率下以约 26 fps 实时生成联合音视频(无需量化),可持续 30 秒且质量稳定。
解读
尚无解读。
原始英文摘要
arXiv:2610.10343v1 Announce Type: new Abstract: Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap. Conventionally, the streaming video literature obtains both capabilities from a chained pipeline. It first distills a bidirectional teacher into a causal student, then into a few-step one, or proceeds in reverse order. Each stage of such a chain fine-tunes the weights the previous one produced, so a later objective can undo an earlier capability. Following the idea of model merging, we show that on a packed audio-video backbone the two capabilities can be acquired in parallel. A causal adapter is trained against the frozen backbone, and an off-the-shelf few-step adapter provides the few-step capability. As the two edit different functional axes, we predict, and then verify, that their weight-update directions are near-orthogonal, without any explicit orthogonality constraint during training. Orthogonal updates should combine without interfering, so parallel composition is a direct sum. The two adapters are simply added at inference, with no joint training, yielding few-step, streaming audio-video whose image quality tracks the bidirectional teacher. Compared to the chained baselines, the composed model matches or beats them on most metrics, making parallel composition a practical approach. The resulting streaming system generates joint audio-video in real time, $\approx$26 fps at $480\times832$ without quantization, and sustains 30 s of continuous generation with stable image quality.