A large web-mined dataset plus variant-consistency reinforcement learning lifts a 7B model to 47% average accuracy on olympiad-style theorem benchmarks, beating similar open models but not top commercial ones.
- Check for adherence to mathematical definitions, theorems, or properties cited in the step
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
A large web-mined dataset plus variant-consistency reinforcement learning lifts a 7B model to 47% average accuracy on olympiad-style theorem benchmarks, beating similar open models but not top commercial ones.