A case study on the effectiveness of llms in verification with proof assistants.arXiv:2508.18587, 2025

Barı¸ s Bayazıt, Yao Li, Xujie Si · 2025 · arXiv 2508.18587

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

representative citing papers

Agentic Proving for Program Verification

cs.AI · 2026-05-22 · unverdicted · novelty 4.0

Agentic Claude reaches 98.8% valid specs, 87.5% implementation certification, and 98.1% end-to-end success on CLEVER, revealing a mismatch between benchmark difficulty and current prover performance.

citing papers explorer

Showing 1 of 1 citing paper.

Agentic Proving for Program Verification cs.AI · 2026-05-22 · unverdicted · none · ref 4
Agentic Claude reaches 98.8% valid specs, 87.5% implementation certification, and 98.1% end-to-end success on CLEVER, revealing a mismatch between benchmark difficulty and current prover performance.

A case study on the effectiveness of llms in verification with proof assistants.arXiv:2508.18587, 2025

fields

years

verdicts

representative citing papers

citing papers explorer