A large web-mined dataset plus variant-consistency reinforcement learning lifts a 7B model to 47% average accuracy on olympiad-style theorem benchmarks, beating similar open models but not top commercial ones.
- If invalid, return False and summarize the critical errors and recommend how to fix the proof/disproof
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
A large web-mined dataset plus variant-consistency reinforcement learning lifts a 7B model to 47% average accuracy on olympiad-style theorem benchmarks, beating similar open models but not top commercial ones.