跳到正文
返回论文列表
cs.AI提交于 已译

TopoGraphRAG-Bench:基于版面布局证据推理的多模态 GraphRAG 评估

TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning

Ruochi Li · Jianzhe Lin · Haoxuan Zhang · Haihua Chen · Junhua Ding · Edward Gehringer · et al.

中文摘要

真实文档中的证据分散于文本、表格、图表与图注,并嵌入复杂的版面布局之中。要在此类文档上回答复杂问题,系统不仅需要检索相关段落,还必须恢复连接异构证据单元的证据拓扑(evidence topology)。现有 GraphRAG 评估仍以文本为中心,而多模态文档 RAG 基准主要考察跨模态检索与生成,并未直接评估对目标证据拓扑的恢复能力。本文提出 TOPOGRAPHRAG-BENCH,一个面向 GraphRAG 中多模态证据推理、基于版面布局的基准,包含 201 份视觉信息丰富的长文档上的 2,024 道问题。问题自底向上由文本、图表与表格证据单元构造,覆盖三种受控拓扑:单跳检索、桥链推理与多源综合。为保证问题保持其目标结构,采用反事实验证以排除捷径、确保模态必要性与证据必要性。在检索、生成与拓扑感知推理指标下,对纯文本 GraphRAG、页面级视觉检索与多模态 GraphRAG 系统进行了评估。多模态 GraphRAG 系统整体表现最强,但在图文证据对齐或多单元组合不完整时仍会出错。纯文本 GraphRAG 在关键依赖来自图表或表格时表现欠佳,页面级视觉检索则缺少恢复拓扑所需的细粒度结构。上述结果推动 GraphRAG 系统从基于文本的实体关系图,转向显式建模文档版面布局、跨模态证据对齐与证据单元的推理角色。代码与数据见 https://richardlrc.github.io/TopoGraphRAG-Bench/。

关键要点

  1. 01问题:现有 GraphRAG 评估以文本为中心,多模态文档 RAG 基准未直接评估证据拓扑恢复能力。
  2. 02方法:基于版面布局构建 TOPOGRAPHRAG-BENCH,含 201 份长文档与 2,024 道问题,覆盖单跳、桥链与多源综合三种受控拓扑,并进行反事实验证。
  3. 03结果:多模态 GraphRAG 整体最强,但在图文对齐或多单元组合不完整时失败;纯文本 GraphRAG 与页面级视觉检索各有不足。
  4. 04局限:多模态 GraphRAG 在跨模态证据对齐与多单元组合场景下仍未解决,需显式建模版面布局与证据单元推理角色。
  5. 05价值:推动 GraphRAG 从文本实体关系图转向同时建模版面布局、跨模态对齐与证据单元角色的方向。

解读

尚无解读。

原始英文摘要

arXiv:2610.09360v1 Announce Type: cross Abstract: Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing GraphRAG evaluations remain largely text-centered, while multimodal document RAG benchmarks assess cross-modal retrieval and generation without directly evaluating recovery of the intended evidence topology. We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents. Questions are constructed bottom-up from text, figure, and table evidence units under three controlled topologies: single-hop retrieval, bridge-chain reasoning, and multi-source synthesis. To ensure that questions preserve their intended structure, we apply counterfactual validation for shortcut resistance, modality necessity, and evidence necessity. We evaluate text-only GraphRAG, page-level visual retrieval, and multimodal GraphRAG systems using retrieval, generation, and topology-aware reasoning metrics. Multimodal GraphRAG systems achieve the strongest overall performance, but still fail when visual-textual evidence alignment or multi-unit composition is incomplete. Text-only GraphRAG struggles when key dependencies are grounded in figures or tables, while page-level visual retrieval lacks the fine-grained structure needed for topology recovery. These findings motivate GraphRAG systems that move beyond text-derived entity relation graphs to explicitly model document layouts, cross-modal evidence alignment, and the reasoning roles of evidence units. Code and data are available at https://richardlrc.github.io/TopoGraphRAG-Bench/.

同方向论文 · cs.AI

查看全部 →