从数字人交互到物理驱动人形机器人技能:面向交互生成器的物理后训练
From Digital Human Interactions to Physics-Based Humanoid Skills: Physics-Grounded Post-Training of Interaction Generators
Kerui Chen · Jianrong Zhang · Kai Lv · Hehe Fan
中文摘要
近期方法在生成两个人形机器人之间的交互方面取得进展,主要依赖基于物理的跟踪策略将参考动作转化为可执行轨迹。然而跟踪能力有限,限制了可成功执行的参考动作范围,降低了数据利用率。此外,即使跟踪成功,也不保证物理上合理的响应或对预期交互的忠实还原。本文提出 DIGHT,一个将数字人交互生成器与人形机器人跟踪策略耦合的协同自适应框架。DIGHT 首先在固定跟踪器下于仿真中执行多个文本条件化的交互候选动作,再依据回放结果构建物理可信偏好,涵盖一般可执行性与交互忠实度。本文将这些信号解耦为物理解耦扩散直接偏好优化(DPO),分别对各准则进行监督而不微分仿真器。为提升可执行性,从跟踪误差、摩擦与漂浮构造偏好对。为提升交互忠实度,提出引入接触保真度等接触反馈,针对接触发生、位置、时长与力幅构建偏好。对齐后的生成器为跟踪器提供参考动作以微调,增强生成与物理执行的兼容性。大量实验表明,该方法不仅提升了生成动作的物理可信度,也实现了仿真中更可靠、更忠实的人形机器人交互。
关键要点
- 01问题:现有跟踪策略对参考动作的跟踪能力有限,导致可执行数据利用率低,且跟踪成功仍未必带来物理合理或忠实于原意的交互响应。
- 02方法:DIGHT 框架协同自适应数字人交互生成器与跟踪策略,通过物理解耦扩散 DPO 将可执行性和交互忠实度偏好分离监督,无需微分仿真器。
- 03方法细节:可执行性偏好基于跟踪误差、摩擦与漂浮构造,交互忠实度偏好则引入仿真器力反馈,针对接触发生、位置、时长与力幅构造。
- 04结果:大量实验显示方法显著提升生成动作的物理可信度,并使仿真中的人形机器人交互更加可靠、忠实。
- 05协同机制:对齐后的生成器为跟踪器提供参考动作进行微调,形成生成与物理执行之间的相互增强闭环。
解读
尚无解读。
原始英文摘要
arXiv:2610.10322v1 Announce Type: new Abstract: Recent methods have made promising progress in generating interactions between two humanoids, largely relying on physics-based tracking policies to convert digital reference motions into executable trajectories. However, limited tracking capabilities restrict the range of reference motions that can be successfully executed, reducing data utilization. Moreover, even successful tracking does not guarantee physically plausible responses or faithful realization of the intended interactions. In this paper, we introduce DIGHT, a co-adaptive framework that couples a Digital human Interaction Generator with a Humanoid Tracking policy. Our DIGHT first executes multiple text-conditioned interaction candidates in simulation using a fixed tracker. It then constructs physics-grounded preferences from the resulting rollouts, covering both general executability and interaction fidelity. Rather than collapsing these signals into a single scalar reward for candidate ranking, we align the pretrained generator using physics-decoupled diffusion direct preference optimization (DPO), preserving criterion-specific supervision without differentiating through the simulator. To improve executability, preference pairs are derived from tracking error, friction, and floating. Additionally, to improve interaction fidelity, we propose to incorporate force feedback from simulator as a measure of contact fidelity and construct preferences over contact occurrence, location, duration, and force magnitude. The aligned generator then supplies reference motions for fine-tuning the tracker, improving compatibility between generation and physical execution. Extensive experiments demonstrate that our approach not only improves the physical plausibility of generated motions but also enables more reliable and faithful humanoid interactions in simulation.