ReverseMath uses answer inversion to generate paired original and reversed math problems with known answers for detecting memorization and improving LLM reasoning via data augmentation.
Evaluating robustness of LLMs to numerical variations in mathematical reasoning
2 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.
2
Pith papers citing it
2
external citations · OpenAlex
years
2026 2representative citing papers
A 13-way text-scrambling benchmark makes open-weight LLMs drop up to 54% average accuracy, and a multi-problem prompt makes their accuracy on the last question decay.
citing papers explorer
-
ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
ReverseMath uses answer inversion to generate paired original and reversed math problems with known answers for detecting memorization and improving LLM reasoning via data augmentation.
-
Robust Reasoning Benchmark
A 13-way text-scrambling benchmark makes open-weight LLMs drop up to 54% average accuracy, and a multi-problem prompt makes their accuracy on the last question decay.