Position Forcing:自条件化三维生成
Position Forcing: Self-Conditioning 3D Generation
Ziheng Ouyang · Zeqiang Lai · Jiarui Chen · Jiangshan Wang · Yuhao Wan · Jingbo Gong · et al.
中文摘要
近期的单阶段三维生成模型普遍采用 VecSet 表示,将三维形状编码为无序的潜在 token 集合。然而,与提供显式位置引导的两阶段方法相比,这类模型必须在去噪过程中隐式推断 token 位置,从而限制了生成质量。研究观察到,尽管缺乏显式的位置条件,VecSet token 仍保留可恢复的空间对应关系。基于此,提出 Position Forcing,一种基于位置的自条件框架。在去噪过程中,Position Forcing 从当前干净潜在估计中恢复 token 位置,按去噪阶段以逐级更细的分辨率进行量化,并将得到的位编码反馈给扩散 Transformer。这种逐级细化的位置反馈为每个去噪阶段提供合适粒度的空间引导,沿由粗到精的轨迹引导形状生成,在无需单独位置生成阶段的情况下显著提升生成质量。实验表明,Position Forcing 在单阶段三维生成方法中取得优异表现,并超越多种具有竞争力的多阶段方法。
关键要点
- 01问题:单阶段 VecSet 三维生成模型缺乏显式位置条件,需在去噪中隐式推断 token 位置,限制生成质量
- 02方法:提出 Position Forcing 自条件框架,去噪时从当前干净潜在恢复 token 位置,按阶段逐级细分辨率量化后反馈给扩散 Transformer
- 03机制:位编码粒度与去噪阶段匹配,实现由粗到精的形状生成轨迹,且无需额外的位置生成阶段
- 04结果:在单阶段三维生成方法中表现优异,并超越多种多阶段竞争方法
- 05洞察:VecSet token 虽无显式位置条件,仍保留可恢复的空间对应关系
解读
尚无解读。
原始英文摘要
arXiv:2610.10342v1 Announce Type: new Abstract: Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens. However, compared with two-stage methods that provide explicit positional guidance, these models must implicitly infer token positions throughout denoising, limiting their generation quality. We observe that, despite the absence of explicit positional conditioning, VecSet tokens retain recoverable spatial correspondences. Building on this observation, we propose Position Forcing, a position-based self-conditioning framework. During denoising, Position Forcing recovers token positions from the current clean latent estimate, quantizes them at progressively finer resolutions according to the denoising stage, and feeds the resulting positional encodings back into the diffusion Transformer. This progressively refined positional feedback provides spatial guidance at a granularity appropriate to each denoising stage, guiding shape generation along a coarse-to-fine trajectory and substantially improving generation quality without a separate position generation stage. Experiments demonstrate that Position Forcing achieves strong performance among single-stage 3D generative methods and outperforms several competitive multi-stage approaches.