On a new Olympiad-style benchmark, outcome-rewarded LLMs reach 80.1% answer accuracy but only 39.7% reasoning correctness, and a step-by-step LLM verifier outperforms holistic judges at spotting flawed reasoning.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning
On a new Olympiad-style benchmark, outcome-rewarded LLMs reach 80.1% answer accuracy but only 39.7% reasoning correctness, and a step-by-step LLM verifier outperforms holistic judges at spotting flawed reasoning.