A two-track Olympiad benchmark with deterministic integer answers and step-by-step proof grading shows frontier LLMs drop sharply versus older math benchmarks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
RIMO: An Easy-to-Evaluate, Hard-to-Solve Olympiad Benchmark for Advanced Mathematical Reasoning
A two-track Olympiad benchmark with deterministic integer answers and step-by-step proof grading shows frontier LLMs drop sharply versus older math benchmarks.