面向灵巧抓取稳定性预测的时序视觉-触觉学习
Temporal Visuo-Tactile Learning for Dexterous Grasp Stability
Ken Nakahara · Aleksei Buvailik · Prokhor Kotov · Roberto Calandra
中文摘要
人类凭借指尖触觉反馈几乎能以完美成功率抓取日常物体,而机器人抓取领域的多数工作仍聚焦于基于视觉的平行夹爪抓取选择。该工作系统研究了高分辨率动态触觉传感对灵巧手抓取稳定性预测及模型引导抓取的作用。为此,研究团队使用配备四枚 Digit 360 触觉传感器的多指机器人手,在 200 个物体上采集了 10,000 次抓取试验,记录每次抓取过程中的外部视觉、本体觉与触觉数据流。基于该数据集,训练了端到端时序多模态模型,从抓取前观测预测抓取后稳定性,并比较了不同感知通道及编码骨干。通过对比实验与受控输入消融表明, 加入触觉特别是高分辨率动态触觉可显著提升抓取稳定性预测性能。进一步将所学预测器作为在线稳定性门控部署于真实机器人,基于视觉-触觉模型的引导重抓使执行后的抓取成功率较无触觉门控提升 10.5 个百分点。结果表明,丰富的指尖传感与能够刻画触觉动态的时序模型可在不显式建模接触力的情况下支持多指手的抓取学习,为从触觉经验到稳定灵巧操作提供了一条可扩展的数据驱动路径。数据集公开地址:https://lasr-lab.github.io/dexterous-grasp-stability/。
关键要点
- 01问题:机器人抓取研究多依赖视觉与平行夹爪,缺乏对灵巧手高分辨率动态触觉的系统评估
- 02方法:在 200 个物体上采集 10,000 次多指抓取试验,融合视觉、本体觉与 Digit 360 触觉,训练端到端时序多模态稳定性预测模型
- 03结果:触觉尤其是高分辨率动态触觉显著提升抓取前观测对抓取后稳定性的预测精度
- 04应用:将预测器作为在线稳定性门控,视觉-触觉引导的重抓使执行后成功率较无触觉门控提高 10.5 个百分点
- 05价值:无需显式接触力建模,仅靠数据驱动即可获得可扩展的灵巧操作路径,并公开相关数据集
解读
尚无解读。
原始英文摘要
arXiv:2610.10283v1 Announce Type: new Abstract: Humans can grasp everyday objects with almost perfect success rates using fingertip tactile feedback, yet much of the robotic grasping literature emphasizes vision-based grasp selection with parallel grippers. In this work, we systematically investigate how high-resolution, dynamic tactile sensing contributes to grasp stability prediction and model-guided grasping in dexterous robotic hands. To this end, we collected a dataset of 10,000 grasp trials across 200 objects using a multi-fingered robotic hand equipped with four Digit 360 tactile sensors, recording external vision, proprioception, and tactile streams throughout each grasp. With this dataset, we trained end-to-end temporal multimodal models to predict post-lift stability from pre-lift grasp observations and compared sensing modalities and encoding backbones. Experimental results and controlled input ablations show that incorporating touch, and particularly high-resolution, dynamic touch, improves grasp stability prediction. Finally, we deployed the learned predictor as an online stability gate on the real robot, where visuo-tactile model-guided regrasping improved the success rate among executed lifts by 10.5 percentage points over a non-tactile gate. These results show how rich fingertip sensing and expressive temporal models that capture the dynamics of touch can support learned grasping with multi-fingered hands without explicit contact or force modeling, providing a scalable data-driven path from tactile experience toward stable dexterous manipulation. The dataset is publicly available at https://lasr-lab.github.io/dexterous-grasp-stability/.