跳到正文
返回论文列表
cs.RO提交于 已译

OpenViTac:面向视触觉策略的统一仿真-真机学习与基准中的统一仿真-真机框架

OpenViTac: Learning and Benchmarking Visuo-Tactile Policies in a Unified Sim-and-Real Framework

Yifan Wu · Qin Li · Nan Min · Guojin Zhong · Haoyu Zhao · Zhiyuan Li · et al.

中文摘要

触觉反馈为具身智能体提供视觉观测之外的物理信息,使其能更可靠地与真实世界交互。然而,尽管视觉-触觉-语言-动作(VTLA, vision-tactile-language-action)策略发展迅速,目前仍缺乏在仿真与真实世界中统一评估触觉机器人操作的基准。为填补这一空白,提出 OpenViTac,一个面向机器人策略的视触觉操作基准,覆盖仿真与真机。OpenViTac 将富含接触的操作任务归纳为四个与触觉相关的能力维度,并提供仿真-真实世界成对设置,以一致地评估 VLA、WAM 与 VTLA(control)策略。基于该基准,研究了不同触觉表征与融合策略对预训练 VLA 模型性能的影响。相应地,提出 OpenVTLA,一个融合最优表征与融合策略的触觉增强框架。此外,利用成对基准设置研究仿真-真机协同训练(sim-real co-training),分析影响跨域策略学习的因素。OpenViTac 为评估并推进视触觉机器人操作提供了统一平台。

关键要点

  1. 01问题:VTLA 策略虽发展迅速,但缺乏在仿真与真实世界之间统一评估触觉机器人操作的基准
  2. 02方法:构建 OpenViTac 基准,涵盖四个触觉相关能力维度,提供仿真-真机成对设置,并提出 OpenVTLA 触觉增强框架以融合最优表征与融合策略
  3. 03结果:在不同触觉表征与融合策略下对 VLA、WAM、VTLA 策略进行一致评估,并通过仿真-真机协同训练分析跨域策略学习
  4. 04局限:摘要未明确给出 OpenViTac 与 OpenVTLA 的量化性能结果或与现有方法的具体对比指标

解读

尚无解读。

原始英文摘要

arXiv:2610.10384v1 Announce Type: new Abstract: Tactile feedback provides embodied agents with physical information beyond visual observations, enabling more reliable interaction with the real world. However, despite the rapid progress of vision-tactile-language-action (VTLA) policies, there remains a lack of unified benchmarks for evaluating tactile-enabled robot manipulation across simulation and the real world. To address this gap, we introduce OpenViTac, a visuo-tactile manipulation benchmark for evaluating robot policies across simulation and the real world. OpenViTac organizes contact-rich manipulation into four tactile-relevant capability dimensions and provides paired simulation-real-world settings for consistent evaluation of VLA, WAM, and VTLA policies. Building upon this benchmark, we investigate how different tactile representations and integration strategies affect the performance of pretrained VLA models. Correspondingly, we introduce OpenVTLA, a tactile augmentation framework that combines the best-performing representation and integration strategy. Furthermore, we leverage the paired benchmark setting to study sim-real co-training and analyze factors affecting cross-domain policy learning. Together, OpenViTac provides a unified platform for evaluating and advancing visuo-tactile robot manipulation.

同方向论文 · cs.RO

查看全部 →