跳到正文
返回论文列表
cs.RO提交于 已译

Agentic RSR:通过场景重建与执行可端到端的机器人策略实现 Real-to-Sim-to-Real

Agentic RSR: Real-to-Sim-to-Real through Scene Reconstruction and Execution-Grounded Robot Policies

Yihan Li · Yating Feng · Shengjiu Sun · Jianing Chen · Hao Ren · Bowen Yang · et al.

中文摘要

针对真实机器人工作空间的仿真需保留与任务相关的交互,基于该仿真开发的策略必须能在真实机器人可获取的观测上运行。然而场景重建与策略开发往往被分开处理。Agentic Real-to-Sim-to-Real(Agentic RSR)框架通过同一操作任务将场景重建、策略开发与真实机器人执行串联起来。给定工作空间视频、任务描述以及已知机器人模型,智能体恢复度量尺度(metric scale),利用视觉反馈迭代优化场景,并在 MuJoCo 中检查与任务相关的交互。随后编码智能体开发可执行策略,从特权物体位姿逐步过渡到视觉观测与随机化仿真。策略可在单次调用中交错多种观测与动作,智能体依据执行反馈选择继续、重试或修改方案。共享的任务级接口将策略与累积经验迁移至真实机器人,由新观测与安全检查指导执行。在涉及两种机器人、共 18 个重建场景的实验中,相对参考深度估计的平均四视角 Depth MAE 为 0.1057 m,平均 Lab ΔE76ₛ 为 11.04,平均灰度 SSIM 为 0.6990。真实机器人实验中,聚合任务成功率可达仿真任务成功率的 80%,表明仿真性能在硬件上得到大幅保留。代码与重建场景数据将公开发布。

关键要点

  1. 01问题:场景重建与策略开发常被分开处理,难以保证仿真交互保真与真实机器人可观测性的统一。
  2. 02方法:Agentic RSR 以同一任务串联重建、策略与执行,智能体迭代优化场景,编码智能体从特权位姿过渡到视觉观测与随机化仿真,并支持观测-动作交错调用。
  3. 03结果:18 个场景平均四视角 Depth MAE 为 0.1057 m,平均 Lab ΔE76ₛ 为 11.04,平均灰度 SSIM 为 0.6990。
  4. 04结果:真实机器人任务成功率达仿真成功率的 80%,仿真性能在硬件上得到显著保留。
  5. 05局限:抽象未提及具体失败模式或边界条件,真实机器人仅覆盖两种机器人平台。

解读

尚无解读。

原始英文摘要

arXiv:2610.10479v1 Announce Type: new Abstract: A simulation of a real robot workspace must preserve task-relevant interactions, while policies developed in it must operate on observations available to the real robot. Yet scene reconstruction and policy development are often treated separately. We present Agentic Real-to-Sim-to-Real (Agentic RSR), a framework that links scene reconstruction, policy development, and real-robot execution through the same manipulation task. Given a workspace video, a task description, and a known robot model, an agent recovers metric scale, iteratively refines the scene using visual feedback, and checks task-relevant interactions in MuJoCo. A coding agent then develops an executable policy, progressing from privileged object poses to visual observations and randomized simulation. The policy can interleave multiple observations and actions within one invocation, while the agent uses execution feedback to continue, retry, or revise its approach. A shared task-level interface carries the policy and accumulated experience to the real robot, where fresh observations and safety checks guide execution. Across 18 reconstructed scenes involving two robots, the mean four-view Depth MAE against reference depth estimates is 0.1057 m, the mean Lab $\Delta E_{76}$ is 11.04, and the mean grayscale SSIM is 0.6990. In real-robot experiments, the aggregate task success rate reaches 80% of the simulation task success rate, indicating substantial retention of simulated performance on hardware. Code and reconstructed scene data will be made publicly available.

同方向论文 · cs.RO

查看全部 →