视频条件生成的联合二维-三维手部运动恢复
Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery
Chen Xu · Yunqi Li · Binbin Huang · Brent Yi · Shenghua Gao · Yi Ma
中文摘要
从视频中恢复准确的三维手部运动仍面临挑战,频繁的遮挡与不完整的视觉观测使得逐帧姿态估计不可靠且时间上不一致。针对这一问题,提出 JoHan,一个统一的生成框架,直接从视频序列恢复手部运动,不依赖中间逐帧姿态预测。该模型从头学习手部二维与三维局部姿态序列的时间演化与跨表征对应关系,联合生成对齐的二维与三维序列。生成的二维轨迹利用二维图像中的直接空间与时间线索,引导后续生成式三维运动重建;所学运动先验促进时间一致性。二维-三维对应关系进一步用于恢复手部相对相机的全局位置与朝向。在多个具有挑战性的基准测试上的实验表明,该方法在局部手部姿态与相机空间重建上的精度与效率均明显提升,所捕手的运动动态显著更平滑,同时保持较高的逐帧姿态精度。
关键要点
- 01问题:从视频恢复三维手部运动受遮挡与观测不全影响,逐帧估计易出现时间不一致
- 02方法:JoHan 统一生成框架,无需中间逐帧姿态预测,联合生成对齐的二维与三维局部手部姿态序列
- 03关键机制:二维轨迹提供空间与时间线索引导三维重建,运动先验保证时间一致性,二维-三维对应恢复相机空间全局位姿
- 04结果:在局部手部姿态与相机空间重建上精度与效率均显著提高,运动平滑度明显优于先前方法
- 05局限:摘要未提及具体失败场景或方法局限
解读
尚无解读。
原始英文摘要
arXiv:2610.10512v1 Announce Type: new Abstract: Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent. To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions. Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence. The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency. Their learned 2D-3D correspondence further enables recovery of the hand's global position and orientation relative to the camera. Extensive experiments on challenging benchmarks demonstrate significantly improved accuracy and speed in local hand-pose and camera-space reconstruction. Notably, our method captures much better hand-motion dynamics, producing significantly smoother motion than previous methods while maintaining high per-frame pose accuracy.