面向自动驾驶的视觉-语言-动作显式几何思维链
Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
Xingtai Gui · Yucheng Zhou · Dongqian Guo · Jiahao Gong · Feiyang Tan · Jianbing Shen
中文摘要
视觉-语言-动作(VLA)模型已成为自动驾驶的一种有前景的范式。然而,现有 VLA 模型仍存在一个根本性错配:驾驶行为需要精确的 3D 几何线索,而视觉-语言的理解与推理基本在 2D 语义空间中进行。本文提出 GeoCoTDrive,一种面向规划任务的显式几何思维链框架,将几何信息以规划为导向进行锚定。GeoCoTDrive 遵循"先用 2D 思考,再用专用 3D 先验驾驶"的范式:首先锚定与决策关键线索对应的 2D 区域,然后在这些区域内从几何基础模型中采样特征,检索局部化的 3D 先验。这些局部几何特征被交错插入自回归上下文,以支持轨迹生成。为监督该过程,本文引入规划相关锚定(planning-relevant grounding)这一新的区域级锚定任务,聚焦于直接影响自车规划决策的局部空间线索,并构建了 PlanningGrounding 数据集,使 VLA 具备面向规划的锚定能力。在多个端到端自动驾驶基准上的实验表明,GeoCoTDrive 持续提升了安全关键场景下的规划性能,验证了显式几何思维链过程对基于 VLA 的规划的有效性。
关键要点
- 01指出 VLA 自动驾驶模型在 2D 语义空间与 3D 几何线索之间的根本性错配问题
- 02提出 GeoCoTDrive 框架,采用"先 2D 思考、后专用 3D 先验驾驶"范式,将局部 3D 几何特征交错注入自回归上下文
- 03引入规划相关锚定任务并构建 PlanningGrounding 数据集,使 VLA 具备面向规划的区域级锚定能力
- 04在多个端到端自动驾驶基准上一致提升安全关键场景下的规划性能
解读
尚无解读。
原始英文摘要
arXiv:2610.10390v1 Announce Type: new Abstract: Vision-language-action~(VLA) models have emerged as a promising paradigm for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a think with 2D first, drive with dedicated 3D priors paradigm. It first grounds 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support the trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.