跳到正文
返回论文列表
cs.AI提交于 已译

Bolzano系统:从专家引导的证明搜索到开放问题自动求解

From Expert-Guided Proof Search to Automated Open-Problem Solving

Adri\'an Z\'ame\v{c}n\'ik · Mat\v{e}j Kripner · Martin Kouteck\'y · Martin Balko · Jan Greb\'ik · Pavel Hub\'a\v{c}ek · et al.

中文摘要

大语言模型正越来越多地参与数学研究,该领域的进展往往依赖于高效的证明搜索、增量改进以及严谨的验证。Bolzano 是一个开源多智能体框架,由并行的证明智能体与一个验证智能体协同组成,并维护一份可读的研究状态记录。在早期阶段,由人类专家挑选问题并提供指引,系统获得了 8 项结果,其证明经领域专家审核通过。受这些案例研究的推动,研究者在约 3800 个开放问题上以无问题特定人工指导的方式运行 Bolzano,攻克了其中约 200 个开放问题。在一项实验中使用了被 STOC(理论计算机科学顶级会议)2026 录用的论文集,并回答了其中论文提出的 4 个问题,相关作者确认了这些结果。

关键要点

  1. 01问题:数学研究中依赖高效的证明搜索、增量改进与严谨验证,亟需能自动化处理开放问题的系统。
  2. 02方法:开源多智能体系统 Bolzano,由并行证明智能体加一个验证智能体组成,并维护人类可读的研究状态。
  3. 03结果:在约 3800 个无人工指导的开放问题中解决了约 200 个;在 STOC 2026 录用论文中回答了 4 个问题并获作者确认。
  4. 04验证:早期专家挑选的案例中,8 项结果的证明经领域专家审核通过。
  5. 05局限:约 200/3800 的求解比例表明系统对大多数开放问题仍无法给出解答,整体能力有限。

解读

尚无解读。

原始英文摘要

arXiv:2610.09769v1 Announce Type: cross Abstract: Large language models are increasingly contributing to mathematical research, where progress often depends on efficient proof search, incremental improvements and careful verification. We describe Bolzano, a multi-agent open-source system that uses parallel prover agents with a verifier agent and maintains a human-readable research state. Initial manual use on expert-selected problems yielded 8 results whose proofs were checked by domain experts. Motivated by these case studies, we ran Bolzano without problem-specific human guidance on about 3,800 open problems extracted from four sets of papers, solving about 200 open problems. One experiment used papers accepted to STOC 2026, a top conference in theoretical computer science. There, we answered four questions raised in the papers, as confirmed by their authors.

同方向论文 · cs.AI

查看全部 →