面向多样地形四足行走的分层强化学习能耗高效步态适应
Energy-Efficient Gait Adaptation via Hierarchical Reinforcement Learning for Quadrupedal Locomotion Across Diverse Terrains
Ammar Issa · Anubhav Singh · Anton Tsaritsin · Sergey Kolyubin
中文摘要
能耗效率是腿式机器人运动控制的关键目标,但要在不同速度范围和地形条件下兼顾低能耗与稳定表现仍具挑战,尤其对步态生成、动作执行和能耗优化紧密耦合的端到端强化学习(RL)策略而言,其对奖励设计十分敏感。研究提出分层强化学习(HRL)框架:由高频策略执行稳定、鲁棒的关节级运动,由低频步态适应模块显式最小化运输成本(CoT)。基于 Isaac 的三阶段训练流程支持零样本仿真到现实(sim-to-real)迁移,并提升跟踪精度、鲁棒性与能效。学习到的分层结构可依据速度自动调整步态:低速采用同侧步(pacing),高速转为小跑(trotting)。仿真结果表明,该方法在广泛指令速度下降低 CoT,同时在平坦、崎不平整和倾斜地形保持鲁棒运动;并在实体 Unitree AlienGo 四足机器人上完成零样本部署,验证实用性。
关键要点
- 01端到端RL耦合步态、执行与能耗,对奖励设计敏感
- 02HRL以高频策略保障关节运动,低频模块最小化CoT
- 03Isaac三阶段训练支持零样本sim-to-real迁移
- 04分层策略自动实现低速pacing到高速trotting
- 05实体Unitree AlienGo部署验证可行性与鲁棒性
解读
尚无解读。
原始英文摘要
arXiv:2610.10297v1 Announce Type: new Abstract: While energy efficiency is a critical objective for legged-robot locomotion control, achieving low energy consumption while maintaining robust performance across different velocity ranges and terrain conditions remains a key challenge. This is particularly true for end-to-end RL policies, where gait generation, motion execution, and energy optimization are tightly coupled, leading to high sensitivity to reward design. In this work, we propose a hierarchical reinforcement learning (HRL) framework that separates a high-frequency policy for stable and robust joint-level motion execution from low-frequency gait adaptation that explicitly minimizes the cost of transport (CoT). The three-stage Isaac-based training procedure enables zero-shot sim-to-real transfer with improved tracking accuracy, robustness, and energy efficiency. The learned hierarchy exhibits automatic speed-dependent gait adaptation, transitioning from pacing at low speeds to trotting at higher speeds. We validate the proposed approach in simulation against representative single-policy and hierarchical locomotion baselines, demonstrating reduced CoT over a broad range of commanded velocities, while maintaining robust locomotion across flat, uneven rough, and inclined terrains. We further demonstrate its practical feasibility through zero-shot deployment on a physical Unitree AlienGo quadruped.