跳到正文
返回论文列表
cs.CV提交于 已译

面向自动驾驶的视觉-语言-动作显式几何思维链

Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving

Xingtai Gui · Yucheng Zhou · Dongqian Guo · Jiahao Gong · Feiyang Tan · Jianbing Shen

中文摘要

视觉-语言-动作(VLA)模型已成为自动驾驶的一种有前景的范式。然而,现有 VLA 模型仍存在一个根本性错配:驾驶行为需要精确的 3D 几何线索,而视觉-语言的理解与推理基本在 2D 语义空间中进行。本文提出 GeoCoTDrive,一种面向规划任务的显式几何思维链框架,将几何信息以规划为导向进行锚定。GeoCoTDrive 遵循"先用 2D 思考,再用专用 3D 先验驾驶"的范式:首先锚定与决策关键线索对应的 2D 区域,然后在这些区域内从几何基础模型中采样特征,检索局部化的 3D 先验。这些局部几何特征被交错插入自回归上下文,以支持轨迹生成。为监督该过程,本文引入规划相关锚定(planning-relevant grounding)这一新的区域级锚定任务,聚焦于直接影响自车规划决策的局部空间线索,并构建了 PlanningGrounding 数据集,使 VLA 具备面向规划的锚定能力。在多个端到端自动驾驶基准上的实验表明,GeoCoTDrive 持续提升了安全关键场景下的规划性能,验证了显式几何思维链过程对基于 VLA 的规划的有效性。

关键要点

  1. 01指出 VLA 自动驾驶模型在 2D 语义空间与 3D 几何线索之间的根本性错配问题
  2. 02提出 GeoCoTDrive 框架,采用"先 2D 思考、后专用 3D 先验驾驶"范式,将局部 3D 几何特征交错注入自回归上下文
  3. 03引入规划相关锚定任务并构建 PlanningGrounding 数据集,使 VLA 具备面向规划的区域级锚定能力
  4. 04在多个端到端自动驾驶基准上一致提升安全关键场景下的规划性能

解读

尚无解读。

原始英文摘要

arXiv:2610.10390v1 Announce Type: new Abstract: Vision-language-action~(VLA) models have emerged as a promising paradigm for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a think with 2D first, drive with dedicated 3D priors paradigm. It first grounds 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support the trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.

同方向论文 · cs.CV

查看全部 →