DeepMath-Creative is a new 179-problem benchmark measuring LLM mathematical creativity through constructive proof and counterexample tasks; the best model, O3 Mini, reaches only about 70% accuracy on basic undergraduate items.
Counterexamples in Real Analysis
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models
DeepMath-Creative is a new 179-problem benchmark measuring LLM mathematical creativity through constructive proof and counterexample tasks; the best model, O3 Mini, reaches only about 70% accuracy on basic undergraduate items.