Pith. sign in

Evaluating robustness of LLMs to numerical variations in mathematical reasoning

2 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.

2 Pith papers citing it
2 external citations · OpenAlex

fields

cs.CL 1 cs.LG 1

years

2026 2

representative citing papers

Robust Reasoning Benchmark

cs.LG · 2026-03-26 · conditional · novelty 6.0 · 2 refs

A 13-way text-scrambling benchmark makes open-weight LLMs drop up to 54% average accuracy, and a multi-problem prompt makes their accuracy on the last question decay.

citing papers explorer

Showing 2 of 2 citing papers.

  • ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation cs.CL · 2026-05-26 · unverdicted · none · ref 31

    ReverseMath uses answer inversion to generate paired original and reversed math problems with known answers for detecting memorization and improving LLM reasoning via data augmentation.

  • Robust Reasoning Benchmark cs.LG · 2026-03-26 · conditional · none · ref 59 · 2 links

    A 13-way text-scrambling benchmark makes open-weight LLMs drop up to 54% average accuracy, and a multi-problem prompt makes their accuracy on the last question decay.