跳到正文
返回论文列表
cs.RO提交于 已译

AirGroundVLN:面向目标导向的空地协同视觉-语言导航的大规模基准

AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation

Zhenxuan Zeng · Qingle Wu · Wei Suo · Maojia Wu · Bairong Zhang · Hangzheng Yu · et al.

中文摘要

目标导向的视觉-语言导航(VLN)要求智能体在无预设路径的情况下,根据自然语言描述定位并抵达目标。空地协同对于同时需要大范围搜索与细粒度定位的任务具有重要价值。然而,目标导向空地协同VLN的系统性研究仍受限于大规模多样化基准的缺乏,以及两个核心挑战:1) 空中与地面视角之间存在显著差异,加之导航过程中可用观测信息持续减少,使得跨平台、跨时间维持空间一致的上下文极为困难;2) 空间观测能力的不对称性使地面感知在局部细节丰富但空间范围有限,空中感知覆盖范围广但局部粒度粗糙,单平台规划的可靠性因此受限。为应对上述局限,本文提出AirGroundVLN基准,涵盖19个Unreal Engine环境中的10,281条导航episode和955个目标实例,设有已见/未见划分及空中可见性协议以支持系统化评估。同时提出可训练参考框架AG-CoNAV,包含两个关键组件:时空锚定协同记忆(SACM)与空中引导的由区域到局部规划(AGRLP)。SACM在空地观测间维护与检索空间一致的历史上下文;AGRLP将区域级空中引导与细粒度地面导航相结合。大量实验验证了AG-CoNAV的有效性,并确立了AirGroundVLN作为未来研究的综合性基准。

关键要点

  1. 01问题:目标导向空地协同VLN缺乏大规模多样化基准,空地视角差异大、有效观测随导航进行而减少,跨平台时空上下文一致性难以维持;单平台空间观测能力不对称,规划可靠性受限
  2. 02方法:构建AirGroundVLN基准,含10,281条导航episode、955个目标实例、19个Unreal Engine环境,提供已见/未见划分与空中可见性协议
  3. 03方法:提出AG-CoNAV框架,包含时空锚定协同记忆SACM与空中引导的区域到局部规划AGRLP,前者维护空地观测间的空间一致历史上下文,后者融合区域级空中引导与细粒度地面导航
  4. 04结果:在AirGroundVLN上的大量实验验证了AG-CoNAV的有效性,并确立该基准作为未来空地协同VLN研究的综合性评测平台
  5. 05价值:为同时需要大范围搜索与精细定位的空地协同导航任务提供了系统性评测基准与方法参考

解读

尚无解读。

原始英文摘要

arXiv:2610.10421v1 Announce Type: new Abstract: Goal-oriented Vision-and-Language Navigation (VLN) requires agents to locate and reach targets described in natural language without prescribed routes. Air--ground collaboration is valuable for tasks requiring both wide-area search and fine-grained localization. However, systematic study of goal-oriented air--ground collaborative VLN remains limited by the lack of large-scale, diverse benchmarks and two core challenges: 1) substantial differences between aerial and ground views, together with useful observations becoming unavailable as navigation proceeds, make it difficult to maintain spatially consistent context across platforms and over time; and 2) asymmetric spatial observability makes ground perception locally detailed but spatially limited and aerial perception broad but locally coarse, limiting the reliability of single-platform planning. To address these limitations, we introduce AirGroundVLN, a benchmark containing 10,281 navigation episodes and 955 target instances across 19 Unreal Engine environments, with seen/unseen splits and an aerial-visibility protocol for systematic evaluation. Alongside the benchmark, we propose AG-CoNAV, a trainable reference framework comprising two key components: Spatiotemporally Anchored Collaborative Memory (SACM) and Aerial-Guided Regional-to-Local Planning (AGRLP). SACM maintains and retrieves spatially consistent historical context across aerial and ground observations. Meanwhile, AGRLP combines regional aerial guidance with fine-grained ground navigation. Extensive experiments demonstrate the effectiveness of AG-CoNAV and establish AirGroundVLN as a comprehensive benchmark for future exploration.

同方向论文 · cs.RO

查看全部 →