A Minimal Agent for Automated Theorem Proving

· 2026 · cs.AI · arXiv 2602.24273

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

open full Pith review browse 3 citing papers arXiv PDF

abstract

We propose a minimal agentic baseline that enables systematic comparison across different AI-based theorem prover architectures. This design implements the core features shared among state-of-the-art systems: iterative proof refinement, library search and context management. We evaluate this agentic approach using qualitatively different benchmarks and compare various frontier language models and design choices. Our results show competitive performance compared to state-of-the-art approaches, while using a significantly simpler architecture and a fraction of their cost. Additionally, we demonstrate consistent advantages of an iterative approach over multiple single-shot generations, especially in terms of sample efficiency and cost effectiveness. The implementation is released open-source as a candidate reference for future research and as an accessible prover for the community.

representative citing papers

AxDafny: Agentic Verified Code Generation in Dafny

cs.AI · 2026-06-30 · unverdicted · novelty 7.0

AxDafny achieves 92.7% verification success on DafnyBench (6.5 points above prior proof-hint baselines) via verifier-guided repair and introduces the LCB-Pro-Dafny benchmark of 250 problems.

Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement

cs.AI · 2026-06-04 · conditional · novelty 7.0

Goedel-Architect introduces blueprint generation and iterative refinement for Lean 4 theorem proving, reaching 99.2% on MiniF2F-test and 75.6% on PutnamBench with DeepSeek-V4-Flash.

Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics

cs.AI · 2026-06-30

citing papers explorer

Showing 3 of 3 citing papers after filters.

AxDafny: Agentic Verified Code Generation in Dafny cs.AI · 2026-06-30 · unverdicted · none · ref 14 · internal anchor
AxDafny achieves 92.7% verification success on DafnyBench (6.5 points above prior proof-hint baselines) via verifier-guided repair and introduces the LCB-Pro-Dafny benchmark of 250 problems.
Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement cs.AI · 2026-06-04 · conditional · none · ref 8 · internal anchor
Goedel-Architect introduces blueprint generation and iterative refinement for Lean 4 theorem proving, reaching 99.2% on MiniF2F-test and 75.6% on PutnamBench with DeepSeek-V4-Flash.
Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics cs.AI · 2026-06-30 · unreviewed · ref 5 · internal anchor

A Minimal Agent for Automated Theorem Proving

fields

years

verdicts

representative citing papers

citing papers explorer