Long-WAM:扩展世界-动作模型的上下文
Long-WAM: Scaling the Context of World-Action Models
Wei Huang · Bohan Zhang · Chenzhi Liu · Isabella Liu · Shuai Yang · Weian Mao · et al.
中文摘要
实时机器人控制需要充足的视觉历史来推断运动与任务进度,但处理这些历史会带来动作延迟。Long-WAM 是一套面向实时控制约束的因果世界-动作模型上下文扩展的模型-系统框架。核心发现是:能否访问历史 ≠ 能否真正利用历史——只有当视频基础模型以自回归(autoregressive, AR)方式预训练时,更长历史才能带来显著收益。具体地,先在机器人与第一人称视频上无动作标签地学习因果预测,再在世界-动作适配阶段保留这种历史到未来的结构。在 RoboCasa GR-1 上,上下文从 0.0 提升到 19.2 秒,成功率由 63.3% 提升至 78.7%;而采用双向预训练初始化则无净增益;在机器人域上进行 AR 预训练进一步提升了 GR-1 与 LIBERO-Long 的峰值成功率。Long-WAM 在 LIBERO-Long、RoboTwin 2.0 与 DOMINO 上均取得对比方法中的最佳结果。流式观测编码、异步执行与硬件特定加速使其可在 RTX 5090、DGX Spark 与 Jetson AGX Thor 上部署且不丢失未来预测能力;在 RTX 5090 上,每个动作块(含未来视频潜变量预测)耗时 107.4 ms。在 Unitree G1 与 YAM 上的实时部署支持动态、长视距操作任务,包括动态堆杯任务 95% 成功率,而 Pi0.5 与 Fast-WAM 在 20 次试验中成功率为 0。作为一种记忆驱动的执行器,Long-WAM 还能与高层规划互补,完成组合任务。
关键要点
- 01问题:实时机器人控制中视觉历史长度与推理延迟难以平衡,长历史未必带来收益
- 02方法:在 AR 视频基础模型上做因果预训练并保留历史到未来结构,再进行世界-动作适配,配合流式编码与异步执行
- 03结果:GR-1 上下文 0.0→19.2 s 成功率 63.3%→78.7%,在 LIBERO-Long、RoboTwin 2.0、DOMINO 上取得对比方法最佳
- 04部署:在 RTX 5090 上每动作块含未来视频预测仅 107.4 ms,可在 G1、YAM 等机器人实时运行
- 05动态长视距表现:动态堆杯任务 95% 成功率,Pi0.5 与 Fast-WAM 20 次均失败,亦可作为高层规划的内存驱动执行器
解读
尚无解读。
原始英文摘要
arXiv:2610.10528v1 Announce Type: new Abstract: Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.