Bolzano系统:从专家引导的证明搜索到开放问题自动求解
From Expert-Guided Proof Search to Automated Open-Problem Solving
Adri\'an Z\'ame\v{c}n\'ik · Mat\v{e}j Kripner · Martin Kouteck\'y · Martin Balko · Jan Greb\'ik · Pavel Hub\'a\v{c}ek · et al.
中文摘要
大语言模型正越来越多地参与数学研究,该领域的进展往往依赖于高效的证明搜索、增量改进以及严谨的验证。Bolzano 是一个开源多智能体框架,由并行的证明智能体与一个验证智能体协同组成,并维护一份可读的研究状态记录。在早期阶段,由人类专家挑选问题并提供指引,系统获得了 8 项结果,其证明经领域专家审核通过。受这些案例研究的推动,研究者在约 3800 个开放问题上以无问题特定人工指导的方式运行 Bolzano,攻克了其中约 200 个开放问题。在一项实验中使用了被 STOC(理论计算机科学顶级会议)2026 录用的论文集,并回答了其中论文提出的 4 个问题,相关作者确认了这些结果。
关键要点
- 01问题:数学研究中依赖高效的证明搜索、增量改进与严谨验证,亟需能自动化处理开放问题的系统。
- 02方法:开源多智能体系统 Bolzano,由并行证明智能体加一个验证智能体组成,并维护人类可读的研究状态。
- 03结果:在约 3800 个无人工指导的开放问题中解决了约 200 个;在 STOC 2026 录用论文中回答了 4 个问题并获作者确认。
- 04验证:早期专家挑选的案例中,8 项结果的证明经领域专家审核通过。
- 05局限:约 200/3800 的求解比例表明系统对大多数开放问题仍无法给出解答,整体能力有限。
解读
尚无解读。
原始英文摘要
arXiv:2610.09769v1 Announce Type: cross Abstract: Large language models are increasingly contributing to mathematical research, where progress often depends on efficient proof search, incremental improvements and careful verification. We describe Bolzano, a multi-agent open-source system that uses parallel prover agents with a verifier agent and maintains a human-readable research state. Initial manual use on expert-selected problems yielded 8 results whose proofs were checked by domain experts. Motivated by these case studies, we ran Bolzano without problem-specific human guidance on about 3,800 open problems extracted from four sets of papers, solving about 200 open problems. One experiment used papers accepted to STOC 2026, a top conference in theoretical computer science. There, we answered four questions raised in the papers, as confirmed by their authors.