TRACEBench shows LRM temporal accuracy falls almost linearly with a controllable Allen-algebra difficulty score, while ~28% of mid-sized model successes are spurious guesses.
Title resolution pending
1 Pith paper cite this work, alongside 5 external citations. Polarity classification is still indexing.
1
Pith paper citing it
5
external citations · external index
fields
cs.SE 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-controlled and Dynamic Test Generation
TRACEBench shows LRM temporal accuracy falls almost linearly with a controllable Allen-algebra difficulty score, while ~28% of mid-sized model successes are spurious guesses.