RobotWorld:面向多类任务与机器人形态的多模态智能体基准评测
RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
Zhiqin Yang · Chenxin Li · Xiaomeng Hu · Yibin Liu · Weidong Huang · Jiankai Sun · et al.
中文摘要
通用智能体已能编写代码、调用工具并完成复杂数字任务,由此引出一个问题:这些能力在多大程度上可以迁移到物理世界。为探究这一问题,论文提出 RobotWorld,一个面向机器人使用的仿真测试床,要求智能体将指令与观测转化为通过机器人接口完成的物理任务执行。该测试床包含 84 个任务,涵盖操作、移动操作、运动、驾驶与空中控制,并设定了显式的交互预算与可执行的成功判定。通过同时分析任务结果与执行轨迹,论文既识别出可迁移的能力,也指出了阻碍任务可靠完成的差距。进一步发现,现有智能体能够构建复杂的感知与控制流程,包括图像分割、相机标定、空间估计与基于动力学的计算。然而,这些能力并不能稳定地组合成成功的行为:智能体在抵达指令位姿后会丢失任务相关物体的状态,无法纠正无效动作,纠偏过晚,或将未完成的任务误判为完成。这种不均衡的迁移在不同模型间表现不同:Astra 在空间类与受限接触类目标上成功率更高,而 Opus 5.5 在连续平衡类与定时交互类目标上成功率更高。RobotWorld 将这些结果与执行行为相关联,既提供了严格的验证平台,也给出了关于剩余能力差距的实证刻画,从而为训练与设计更可靠的物理世界智能体确立了具体目标。
关键要点
- 01问题:通用数字任务能力能否迁移到物理世界,智能体在真实机器人执行中的可靠性尚未被系统评测
- 02方法:构建 RobotWorld 仿真测试床,含 84 个任务,覆盖操作、移动操作、运动、驾驶与空中控制,设定交互预算与可执行成功判定,并结合任务结果与执行轨迹分析
- 03结果:智能体可构建图像分割、相机标定、空间估计与动力学计算等流程,但难以将其稳定组合为成功行为,存在状态丢失、纠偏失败、过早终止等问题
- 04差异:模型间能力迁移不均衡,Astra 在空间与受限接触任务上更强,Opus 5.5 在连续平衡与定时交互任务上更强
- 05意义:建立物理世界智能体能力差距的实证基线,为训练与设计更可靠的智能体提供具体目标
解读
尚无解读。
原始英文摘要
arXiv:2610.10409v1 Announce Type: new Abstract: General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.