跳到正文
返回论文列表
cs.AI提交于 已译

GAGR-Lab:联合空间几何与分析函数推理评估

GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning

Jingyao Zhang · Yun Li · Lu Han

中文摘要

联合空间几何与分析函数推理要求将感知到的空间构型转化为符号函数,使其执行曲线满足几何约束。提出 GAGR-Lab 框架,通过笛卡尔游戏场景、显式函数语义与权威 Rust 轨迹执行来度量该能力,并区分空间感知、度量基准化、几何关系、函数解释、函数构造与约束合成六个子能力。框架配置四档场景难度预设与一个前瞻性的 24 格诊断设计,仅报告实际评估的子集。一次有限预实验以一款托管模型 Llama 3.2 11B Vision Instruct 配合两组 API 凭证作为执行副本,生成 72 局平衡游戏、432 次尝试,收到 429 条有效提供者响应,无任何目标命中;探索性普通函数提示变体同样未命中,结构化定位接口未产出可评分输出。特权分析搜索控制在 300 个生成场景中的 600 个方向用例上独立成功,具备完全可复现性,并通过 1200 次垂直反射或平移校验。框架将服务可靠性、符号合规性与几何成功加以区分,保留精确的模型可见输入与实际路径。分阶段协议涵盖诊断校准、留出复现、多模型对比与配对鲁棒性测试。贡献是一个带已执行预实验与明确前瞻研究计划的可运行研究框架,完整难度矩阵与模型对比结果尚未测试。

关键要点

  1. 01定义联合空间几何与分析函数推理任务,需将空间构型映射为满足几何约束的符号函数
  2. 02提出 GAGR-Lab 框架,覆盖六项子能力评估与四档场景难度预设,配套 24 格诊断设计
  3. 03Llama 3.2 11B Vision Instruct 预实验 432 次尝试零命中,服务可靠性、符号合规性与几何成功分层统计
  4. 04特权分析搜索控制通过 600 方向用例与 1200 次反射/平移校验,验证评估管线可复现性
  5. 05完整难度矩阵与多模型对比仍属前瞻计划,尚未执行

解读

尚无解读。

原始英文摘要

arXiv:2610.10201v1 Announce Type: cross Abstract: Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and authoritative Rust trajectory execution. It distinguishes spatial perception, metric grounding, geometric relations, function interpretation, function construction, and constrained synthesis. We specify four configurable scene-difficulty presets and a prospective 24-cell diagnostic design, while reporting only the subset actually evaluated. A bounded pilot of one hosted model (Llama 3.2 11B Vision Instruct) using two API credentials as execution replicas yields 72 balanced games with 432 attempts, 429 valid provider responses, and no target hits; exploratory ordinary-function prompt variants also fail to hit, while the structured localization interface yields no scoreable outputs. A privileged analytic search control independently succeeds on 600 directional cases from 300 generated scenes, with exact repeatability and 1,200 successful vertical-reflection or translation checks. The framework separates serving reliability, symbolic compliance, and geometric success, and preserves exact model-visible inputs and realized paths. A staged protocol outlines diagnostic calibration, held-out replication, multi-model comparison, and paired robustness tests. The contribution is an operational research framework with an executed pilot and a clearly identified prospective study plan; the full difficulty matrix and comparative model results remain untested.

同方向论文 · cs.AI

查看全部 →