RFPO:面向具身控制的整流流策略优化
RFPO: Rectified Flow Policy Optimization for Embodied Control
Ting Huang · Lisiyu Pan · Haoyu Wang · Zeyu Zhang · Siyuan Qian · Yanjun Li · et al.
中文摘要
基于流的策略为连续机器人控制提供了富有表达力的框架,但其迭代 ODE(常微分方程)积分带来显著的推理成本。简单削减积分步数会严重损害控制效果,因为在全步执行下优化的策略并未被显式约束在粗粒度数值积分下仍保持可靠。这一不匹配被称为少步离散化差距(few-step discretization gap)。为此,提出 RFPO,一个面向可靠少步执行的流策略优化框架。奖励感知的在线 Reflow 在 on-policy 学习过程中矫正学生策略诱导的传输路径,使其对粗粒度积分更鲁棒。一个冻结的高斯 PPO 控制器在全步和中间积分预算下提供互补的动作空间监督,而部署时仍是仅以单步 Euler 积分执行的一个流学生策略。在 Unitree Go2、Boston Dynamics Spot、Unitree H1 和 Unitree G1 上,RFPO 在一步执行下始终保持全步控制性能,无论零初始化还是随机初始化,一步回报均保持在对应 64 步回报的 2.4% 以内。在 Unitree Go2 上,一步执行保留了 64 步奖励的 98.5%,同时将板载平均推理延迟由 4.39 ms 降至 0.08 ms,实现 54.9 倍加速。真实机器人实验进一步验证了一步步态的稳定性。代码:https://github.com/AIGeeksGroup/RFPO。项目页:https://aigeeksgroup.github.io/RFPO。
关键要点
- 01问题:流策略 ODE 积分推理开销大,粗粒度积分会引入少步离散化差距,损害控制性能
- 02方法:RFPO 框架,结合奖励感知在线 Reflow 矫正传输路径,并用冻结高斯 PPO 教师提供动作空间监督
- 03结果:在 Go2、Spot、H1、G1 上一步执行与 64 步回报差距在 2.4% 以内,Go2 推理延迟从 4.39 ms 降至 0.08 ms(54.9× 加速)
- 04结果:Go2 一步执行保留 64 步奖励的 98.5%,真实机器人实验验证了一步步态稳定性
- 05局限:摘要未提及方法在其他任务或非腿式机器人上的泛化能力以及长期部署下的累积误差问题
解读
尚无解读。
原始英文摘要
arXiv:2610.10453v1 Announce Type: new Abstract: Flow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost. Naively reducing the integration budget can severely degrade control, since policies optimized under full-step execution are not explicitly constrained to remain reliable under coarse numerical integration. We refer to this mismatch as the few-step discretization gap. To address this problem, we introduce RFPO, a flow-policy optimization framework for reliable few-step execution. Reward-aware online Reflow rectifies student-induced transport paths during on-policy learning, making the resulting policy more robust to coarse integration. A frozen Gaussian PPO controller supplies complementary action-space supervision at full and intermediate integration budgets, while the deployed policy remains a single flow student executed with one Euler step. Across Unitree Go2, Boston Dynamics Spot, Unitree H1, and Unitree G1, RFPO consistently preserves full-step control performance under one-step execution, with one-step returns remaining within 2.4% of their corresponding 64-step values across both zero and random initialization. On Unitree Go2, one-step execution retains 98.5% of the 64-step reward while reducing onboard mean inference latency from 4.39 ms to 0.08 ms, yielding a 54.9x speedup. Real-robot experiments further validate stable one-step locomotion. Code: https://github.com/AIGeeksGroup/RFPO. Website: https://aigeeksgroup.github.io/RFPO.