HuMBLE:基于人类运动驱动的具身运动行为学习
HuMBLE: Human Motion-Driven Behavior Learning for Embodied Locomotion
Mike Zhang · Dongho Kang · Kevin Bergamin · Nicola Burger · Robin Deits · Jonathan Foster · et al.
中文摘要
尽管近期仿人机器人运动控制取得了进展,面向指令跟踪与鲁棒性优化的控制器往往生成机械化的步态,而依赖人类运动数据的控制器又难以泛化到数据分布之外的指令。该工作提出一种学习框架,用于平衡这些相互竞争的目标,从人类运动数据合成实时可操控、鲁棒且仿生的运动控制策略。利用内部策划的涵盖多种速度与方向的运动数据集,首先通过师生蒸馏过程学习一个自然运动先验策略。具体方法是,用强化学习(RL)训练一个全身参考条件策略,再将其蒸馏为一个仅以本体感知和平面躯干速度操控命令为主,为速度驾驶命令为条件的轻量策略。接下来,用多任务强化学习对先验策略进行微调,以扩展指令覆盖范围与超出数据分布之外的鲁棒性,组合一个跟踪任意指令的目标条件任务和一个将人类数据作为显式风格正则化器的参考引导任务。该框架在三种仿人机器人上验证:Boston Dynamics Atlas R1、Atlas D1 与 Unitree G1。实验结果表明,该策略在真实场景下表现鲁棒,包括室内外用户直接操控运动,以及作为分层控制栈中的运动层集成。与不使用人类数据训练的 Tabula Rasa 强化学习策略的对比基准测试与消融研究证实,该框架可产出一种轻量、可部署的策略,能从一个驾驶命令重建协调的全身行为,在保持人类步态特征的同时仍具鲁棒性与完全可操控性。
关键要点
- 01问题:面向指令跟踪与鲁棒性的仿人控制器产生机械化步态,而基于人类数据的控制器难以泛化到数据分布外的指令,两类方法各有取舍
- 02方法:利用自建多速度、多方向运动数据集,通过师生蒸馏从全身参考条件 RL 策略蒸馏出仅依赖本体感知与平面躯干速度命令的轻量先验策略,再用多任务 RL 组合目标条件跟踪任务与风格正则化任务进行微调
- 03结果:在 Atlas R1、Atlas D1 与 Unitree G1 三种机器人上实现室内外用户直接操控与分层控制栈集成,具备实时可操控性与鲁棒性
- 04对比与消融:对比不使用人类数据的 Tabula Rasa RL 策略及消融研究,证实人类数据风格正则化对保持步态自然度的必要性
- 05局限:先验策略依赖自建运动数据集覆盖的速度与方向范围,超出分布的鲁棒性依赖多任务 RL 微调扩展
解读
尚无解读。
原始英文摘要
arXiv:2610.10489v1 Announce Type: new Abstract: Despite recent advances in humanoid locomotion, controllers optimized for command tracking and robustness tend to produce mechanical gaits, whereas controllers tied to human motion data often fail to generalize to commands outside the data distribution. This work introduces a learning framework that balances these competing objectives to synthesize real-time steerable, robust, and biomimetic locomotion policies from human data. Using an in-house curated locomotion dataset covering diverse speeds and directions, we first learn a natural locomotion prior policy through a teacher-student distillation process. Specifically, we train a full-body reference-conditioned policy with Reinforcement Learning (RL), then distill it into a lightweight prior policy conditioned solely on proprioception and a planar torso-velocity steering command. Next, we fine-tune the prior policy with multi-task RL to expand command coverage and robustness beyond the data distribution, pairing a goal-conditioned task that tracks arbitrary commands with a reference-guided task that tracks the human data as an explicit style regularizer. We validate our framework on three humanoid robots: the Boston Dynamics Atlas R1, Atlas D1, and Unitree G1. Experimental results demonstrate robust performance across real-world scenarios, including direct user-controlled locomotion in indoor and outdoor environments, and integration as the locomotion layer within hierarchical control stacks. Benchmarks against Tabula Rasa RL policies trained without human data and ablation studies confirm that our framework yields a lightweight, deployable policy that reconstructs coordinated whole-body behavior from a steering command, retaining the human gait characteristics while remaining robust and fully steerable.