A financial-agent benchmark with 220 open-ended queries and 11,543 source-attributed rubrics finds that the tool harness shapes performance more than the model alone, and that the authors' in-house system leads at 56%.
Fi- nanceReasoning: Benchmarking financial numerical reasoning more credible, comprehensive and challenging
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents
A financial-agent benchmark with 220 open-ended queries and 11,543 source-attributed rubrics finds that the tool harness shapes performance more than the model alone, and that the authors' in-house system leads at 56%.