REVIEW 11 cited by
A Diverse Corpus for Evaluating and Developing English Math Word Problem Solvers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present ASDiv (Academia Sinica Diverse MWP Dataset), a diverse (in terms of both language patterns and problem types) English math word problem (MWP) corpus for evaluating the capability of various MWP solvers. Existing MWP corpora for studying AI progress remain limited either in language usage patterns or in problem types. We thus present a new English MWP corpus with 2,305 MWPs that cover more text patterns and most problem types taught in elementary school. Each MWP is annotated with its problem type and grade level (for indicating the level of difficulty). Furthermore, we propose a metric to measure the lexicon usage diversity of a given MWP corpus, and demonstrate that ASDiv is more diverse than existing corpora. Experiments show that our proposed corpus reflects the true capability of MWP solvers more faithfully.
Forward citations
Cited by 11 Pith papers
-
FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models
FlowBlock overlaps adjacent blocks of self-correcting diffusion LLMs, achieving up to 2.95x and 4.01x higher tokens/sec over serial baselines with up to 77.1% lower latency and matched or better accuracy.
-
Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia
Targeted perturbation of reward-anticipatory units in VLMs induces anhedonia-like effort avoidance and clinical-scale score drops without impairing baseline task competence.
-
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis
Shortcut neuron patching suppresses benchmark-contamination shortcuts in LLMs and yields evaluation scores that strongly correlate with the external MixEval benchmark.
-
MultiLingPoT: Enhancing Mathematical Reasoning with Multilingual Program Fine-tuning
Training one model on Program-of-Thought solutions in four programming languages raises math accuracy for each language, and answer mixing outperforms single-language augmented training by up to about 6 percentage points.
-
Efficiently Scaling LLM Reasoning with Certaindex
Certaindex measures when an LLM's intermediate answers stop changing, enabling early exit and dynamic token allocation that cuts token usage by up to 50% with no accuracy drop in tested workloads.
-
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
The thesis demonstrates that combining implicit 3D scene representations with LLM-based reasoning, using text as an interface, yields strong performance on robotic perception and spatial language tasks.
-
Reason from Future: Reverse Thought Chain Enhances LLM Reasoning
A prompting method that alternates backward and forward reasoning improves small LLM accuracy on math and search tasks and reduces the number of visited search states.
-
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey
A comprehensive review that categorizes methods for shortening and adaptively triggering chain-of-thought reasoning in large language models.
-
Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions
LLM robustness research is organized into adversarial robustness, out-of-distribution robustness, and evaluation, with an accompanying GitHub collection of papers.
-
A Survey on Large Language Models with some Insights on their Capabilities and Limitations
A broad survey of LLM methods and applications, plus an empirical section on how code-rich pretraining may influence chain-of-thought reasoning, the details of which are not visible in the supplied text.
-
A Survey on Large Language Models for Mathematical Reasoning
Recent advances in LLM mathematical reasoning are organized into comprehension and generation phases, covering methods from prompting to test-time scaling.
Discussion (0). Continue with ORCID to comment.