跳到正文
返回论文列表
cs.RO提交于 已译

RoboPrompt:基于稀疏人类输入的直觉式机器人策略引导

RoboPrompt: Intuitive Robot Policy Steering with Sparse Human Input

Yanwen Zou · Chenyang Shi · Guoxuan Xu · Wenye Yu · Wendi Chen · Ye Pan · et al.

中文摘要

端到端的模仿学习机器人策略受限于数据多样性不足,在真实场景中的零样本部署仍然困难。共享自主方法通过遥操作实现人工纠错,然而专用硬件和操作员培训阻碍了其大规模部署。其他方法将人类引导作为策略的额外输入,通常需要修改网络结构并针对可引导性进行额外训练,从而限制了其跨策略的通用性。本文提出 RoboPrompt,一种通用、轻量的机器人策略引导系统,支持用户通过稀疏直觉输入(手绘轨迹、目标点、粗略方向指令)引导策略行为。RoboPrompt 将意图转换与底层策略解耦:一个可复用模块将引导信号转换为动作草案,再经由基策略的扩散(diffusion)或流匹配(flow-matching)动力学进行精修。通过在噪声空间控制动作生成,RoboPrompt 在不修改基策略架构或针对可引导性进行微调的前提下,平衡了人类意图与策略先验。实验表明,该方法可对 Diffusion Policy、π₀.₅ 和 FastWAM 实现有效引导,并可结合引导回滚进行在线策略改进(DAgger: Dataset Aggregation)。经过 2-3 轮迭代后,π₀.₅ 在三项任务上的平均成功率提升 15.5%,在 Insert Bread 任务上,三种策略(Diffusion Policy、π₀.₅、FastWAM)的平均成功率提升 21.3%,平均人工介入次数分别下降 44.0%(2.86 降至 1.60)和 81.9%(2.60 降至 0.47)。

关键要点

  1. 01数据多样性不足导致模仿学习策略零样本真实部署困难,而现有共享自主和人类引导方法各自受限于硬件依赖或架构侵入。
  2. 02RoboPrompt 通过可复用模块将稀疏人类引导(轨迹、目标点、方向)转为动作草案,再利用基策略自身的扩散或流匹配动力学在噪声空间进行精修。
  3. 03该方法无需修改基策略架构或针对可引导性微调,可即插即用地引导 Diffusion Policy、π₀.₅、FastWAM 等多种策略。
  4. 04结合 DAgger 在线策略改进,2-3 轮迭代后,π₀.₅ 在三项任务上成功率提升 15.5%,三种策略在 Insert Bread 上成功率提升 21.3%。
  5. 05平均人工介入次数在 Insert Bread 上由 2.60 降至 0.47(降幅 81.9%),显著降低部署过程中的人类干预负担。

解读

尚无解读。

原始英文摘要

arXiv:2610.10534v1 Announce Type: new Abstract: End-to-end robot policies trained through imitation learning remain constrained by limited data diversity, making reliable zero-shot deployment in real-world settings challenging. Shared-autonomy methods enable human correction through teleoperation, but specialized hardware and operator training hinder deployment at scale. Other approaches incorporate human guidance as additional policy inputs, often requiring architectural changes and dedicated training for steerability, which limits their applicability across policies. We present RoboPrompt, a general-purpose, lightweight robot policy steering system that enables users to guide policy behavior through intuitive, sparse inputs, including drawn traces, target points, and coarse directional instructions. RoboPrompt decouples human-intention translation from the underlying policy: a reusable module converts human guidance into action drafts, which are refined through the diffusion or flow-matching dynamics of the base policy. By controlling action generation in noise space, RoboPrompt balances human intent with the policy prior without modifying the base policy architecture or fine-tuning it for steerability. Experiments demonstrate effective steering across Diffusion Policy, $\pi_{0.5}$, and FastWAM. We further use steered rollouts for online policy improvement through DAgger. After 2-3 rounds of iteration, average success rates increase by 15.5\% for $\pi_{0.5}$ across three tasks and by 21.3\% across three policies(Diffusion Policy, $\pi_{0.5}$, FastWAM) on the Insert Bread task, while average human intervention counts decrease by 44.0\% (2.86 to 1.60) and 81.9\% (2.60 to 0.47), respectively.

同方向论文 · cs.RO

查看全部 →