A re-annotated and expanded financial reasoning benchmark with a Python function library, where the best model (OpenAI o1 with program-of-thought) reaches 89.1% on 238 hard problems.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
A re-annotated and expanded financial reasoning benchmark with a Python function library, where the best model (OpenAI o1 with program-of-thought) reaches 89.1% on 238 hard problems.