RoboQuest:搜索、检验与测试的通用物理智能体
RoboQuest: Generalist Physical Agents that Search, Inspect and Test
Liu Renhang · Navonil Majumder · Tej Deep Pala · Soujanya Poria
中文摘要
多模态(multimodal)基础模型的进展,使其能够胜任多种操作任务。然而,在陌生环境中成功行动,往往要求智能体在观测缺失相关信息时通过交互主动寻找:判断相关物体位置、检查未观测属性,或发现陌生工具的效果。RoboQuest 是面向目标导向具身探索的基准,要求智能体通过物理交互主动获取任务相关信息,利用所得证据调整后续行动,并自主决定何时执行任务。基准包含 10 项移动操作任务,覆盖搜索、基于操作的检验和交互式测试三类不确定性。研究通过统一的视觉—运动接口评估 5 个前沿多模态智能体,并评估在公开的全回合演示上微调的 $\pi_{0.5}$ 策略。最佳智能体仅在 23% 的回合中成功,微调策略几乎无法成功。单独测试任务所需执行技能并直接提供隐藏信息时,智能体能完成大多数动作;失败分析表明,仅少数失败源于执行。智能体常在尚未观察到完成任务所需证据时就过早停止探索,也很少预防或修复探索造成的干扰。此外,大多数模型仍难以通过试错学习。
关键要点
- 01提出目标导向具身探索基准 RoboQuest,含 10 项移动操作任务
- 02聚焦搜索、操作检验与交互式测试三类不确定性
- 03最佳多模态智能体的回合成功率仅为 23%
- 04多数失败源于过早停止探索,而非动作执行能力不足
- 05智能体难以预防或修复探索干扰,试错学习仍然困难
解读
尚无解读。
原始英文摘要
arXiv:2610.10388v1 Announce Type: new Abstract: Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a $\pi_{0.5}$ policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23\% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution. Our failure analysis further finds that the agents often stop exploring too early as they make decisions before observing the required evidence for task completion. We also find that agents rarely prevent or repair the disturbances caused by their exploration. Moreover, learning by trial and error remains difficult for most models.