REVIEW 66 cited by
Large Language Models for Mathematical Reasoning: Progresses and Challenges
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Mathematical reasoning serves as a cornerstone for assessing the fundamental cognitive capabilities of human intelligence. In recent times, there has been a notable surge in the development of Large Language Models (LLMs) geared towards the automated resolution of mathematical problems. However, the landscape of mathematical problem types is vast and varied, with LLM-oriented techniques undergoing evaluation across diverse datasets and settings. This diversity makes it challenging to discern the true advancements and obstacles within this burgeoning field. This survey endeavors to address four pivotal dimensions: i) a comprehensive exploration of the various mathematical problems and their corresponding datasets that have been investigated; ii) an examination of the spectrum of LLM-oriented techniques that have been proposed for mathematical problem-solving; iii) an overview of factors and concerns affecting LLMs in solving math; and iv) an elucidation of the persisting challenges within this domain. To the best of our knowledge, this survey stands as one of the first extensive examinations of the landscape of LLMs in the realm of mathematics, providing a holistic perspective on the current state, accomplishments, and future challenges in this rapidly evolving field.
Forward citations
Showing 60 of 66 Pith papers that cite this
-
Superloop Equations and Minimal Surfaces I: Confining minimal surface in $4D, N=1$ SYM
A geometrically constructed surface-area phase is proven to dress any solution of the finite-N N=1 SYM superloop hierarchy and produces a rectangular Wilson phase exp(-iσLT) with arbitrary positive σ.
-
AL-Bench: A Benchmark for Automatic Logging
A new benchmark with static and runtime evaluation shows state-of-the-art automatic logging tools produce many uncompilable or semantically misaligned log statements.
-
GaussMark: A Practical Approach for Structural Watermarking of Language Models
GaussMark embeds a detectable watermark by adding per-generation Gaussian noise to one weight matrix and detecting gradient alignment with that noise.
-
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
A one-step semiparametric estimator using pairwise comparison signals as control variates achieves the efficiency bound for estimating LLM accuracy on math benchmarks.
-
Adaptive Information Control for Search-Augmented LLM Reasoning
DeepControl uses information-utility signals to control when search-augmented reasoning agents stop retrieving and how much evidence they expand, improving QA accuracy across seven benchmarks and two model sizes.
-
One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents
Repository-level issue localization can be done by a single jump-to-definition tool trained with reinforcement learning, achieving strong results on SWE-bench despite using only open-weights models.
-
SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials
A new multimodal benchmark for strength of materials shows current foundation models solve at most 56.6% of problems, and expert-written diagram descriptions help more than images.
-
Arrows of Math Reasoning Data Synthesis for Large Language Models: Diversity, Complexity and Correctness
A program-assisted pipeline generates 12.3 million math problem-solution pairs with execution-based verification, and fine-tuning on a 50k sample improves model scores on GSM8K, MATH, Minerva, and SVAMP.
-
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
A generative multimodal process reward model that produces step-level critiques and corrections improves average math accuracy for six multimodal LLMs by 2.9 to 5.9 points under a refinement-based Best-of-N strategy.
-
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
A new wargame-based benchmark finds large language models score far below human experts on strategic reasoning across situation awareness, opponent modeling, and policy generation.
-
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
Reasoning fine-tuning makes LLMs more accurate on answerable problems but worse at abstaining on unanswerable ones, across a new 20-dataset benchmark.
-
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
Safe uses step-level formal verification in Lean 4, aggregated by a small LSTM and combined with process reward scores, to improve best-of-n accuracy for LLM mathematical reasoning.
-
Structured Pruning for Diverse Best-of-N Reasoning Optimization
SPRINT learns to select which attention heads to prune per question, improving Pass@N over random head selection and multinomial sampling on MATH500 and GSM8K.
-
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
Comparative words in prompts can shift LLM answers toward the framed direction in simple arithmetic comparisons, with demographic terms amplifying the effect.
-
Learning to Insert [PAUSE] Tokens for Better Reasoning
A likelihood-based [PAUSE] token insertion method for fine-tuning shows small gains on GSM8K and MBPP, but the AQUA-RAT result is unreliable because the test set contains training samples.
-
Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
LLM reasoning can be scored separately for knowledge and step-by-step information gain, and doing so shows SFT and RL affect these two capacities differently across medicine and math.
-
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
PCPO selects preference pairs by combining correct-answer status with token-level probability consistency, then trains with a weighted DPO+NLL loss, yielding small and partly inconsistent gains over outcome-only metho...
-
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis
A multilingual critique-data training paradigm with a knowledge-unit reward improves LLM cultural alignment on several benchmarks, but its headline benchmark is evaluated with the same LLM-judged metric used to select...
-
Can reasoning models comprehend mathematical problems in Chinese ancient texts? An empirical study based on data from Suanjing Shishu
Reasoning LLMs partially solve classical Chinese math problems from Suanjing Shishu, reaching up to 70% accuracy with original solution methods provided, but lag behind their modern-math performance.
-
Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts
Transformers outperformed young children and matched their error profile on a geometry odd-one-out task, while vision-language models underperformed vision-only models.
-
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
A tree-search theorem prover trained on synthetic proof-state exploration data reaches 60.74% Pass@1 on MiniF2F and 21.18% on ProofNet using an adaptive beam size.
-
Time-R1: Towards Comprehensive Temporal Reasoning in LLMs
A 3B model trained by staged reinforcement learning with rule-based rewards claims to outperform 671B models on temporal prediction and generation, though test-set checkpoint selection and synthetic training data weak...
-
CellVerse: Do Large Language Models Really Understand Cell Biology?
CellVerse evaluates 14 LLMs on language-formatted single-cell tasks and finds that generalist models perform poorly, with drug response prediction near random guessing.
-
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges
CipherBank evaluates 16 LLMs on 2,358 known-plaintext decryption tasks across 9 ciphers; the best model (Claude-3.5) achieves 45.1% accuracy, and o1 reaches 40.6%.
-
Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask
With context-rich prompts that include CWE hints and marked potential vulnerability sites, LLM detectors beat random baselines, but the evaluation may leak the answer.
-
Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration
A symbolic geometry engine generates step-by-step training data and verifies MLLM reasoning steps, improving accuracy on geometry benchmarks.
-
Optimizing Temperature for Language Models with Multi-Sample Inference
Selecting the temperature at the entropy turning point of a language model's generated text yields near-optimal multi-sample inference accuracy without labeled validation data.
-
Adversarial Reasoning at Jailbreaking Time
A loss-guided 'reason, verify, search' loop with three LLM modules surpasses prior semantic jailbreak methods on several defended models.
-
Language Models Use Trigonometry to Do Addition
Three LLMs represent two-digit numbers as generalized helices and appear to compute addition by combining these helices into an answer helix.
-
UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models
UGMathBench provides a 5,062-problem dynamic benchmark for undergraduate math reasoning and shows that leading LLMs solve all versions of only about half the problems.
-
An Automatic Graph Construction Framework based on Large Language Models for Recommendation
AutoGraph uses LLM semantic vectors, residual vector quantization, and metapath GAT propagation to construct an automatic graph that improves recommendation across four backbones and three datasets.
-
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
A prompt-only LLM framework that generates and executes Python code answers 90 to 93 percent of 100 GTFS transit-data queries without fine-tuning.
-
Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement
LLM numeracy failures can be usefully organized as separate representational and procedural grounding problems, with procedural arithmetic improving more than number understanding when models are asked to reason.
-
Assessing the Benefits of Combining Advanced Deep Learning Techniques for Post-Disaster Building Damage Assessment from UAV Imagery
A hybrid CV+LVLM pipeline improves post-disaster building damage counting over single models in some configurations, but fails in others and shows low absolute accuracy.
-
STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA
Compressing multi-hop search trajectories into per-answer evidence cards improves final answer selection over raw-trajectory or string-only comparison on four multi-hop QA benchmarks.
-
Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
LLM math accuracy varies across equivalent problem representations, and executable-reasoning scaffolding redistributes rather than removes the errors.
-
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
LLM reasoning benchmark scores vary substantially across repeated runs under the same model, strategy, and task, so single-run evaluation can misrank systems.
-
LLMs-guided adaptive compensator: Bringing Adaptivity to Automatic Control Systems with Large Language Models
An LLM-generated adaptive compensator, refined through iterative prompting, outperformed classical adaptive controllers on soft and humanoid robots in simulation and prototype tests.
-
Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding
DeepSeek-R1-distilled models show higher multi-document QA accuracy than their base counterparts and flatter position-bias curves, especially with 50-80 documents.
-
Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs
Across ten LLMs, masking the final answer inside a complete reasoning chain causes a 26.9-point accuracy drop, evidence that models anchor to answers, not reasoning templates.
-
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
PrefBERT, a 150M-parameter reward model trained on human quality ratings, outperforms ROUGE-L and BERTScore as a GRPO reward signal for open-ended long-form generation.
-
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
A new benchmark (POBs) reveals that LLMs lean progressive-collectivist, that test-time compute offers limited gains in neutrality or consistency, and that newer model versions often become more biased and less consistent.
-
Towards General Continuous Memory for Vision-Language Models
A vision-language model can act as its own continuous memory encoder, compressing external multimodal knowledge into eight embeddings that improve reasoning when prepended to the frozen model.
-
Lightweight Latent Verifiers for Efficient Meta-Generation Strategies
LiLaVe, an XGBoost verifier trained on hidden states of a base LLM, predicts answer correctness with AUC comparable to large LLM-based verifiers, and enables conditional majority voting and self-correction that improv...
-
Schema-Guided Scene-Graph Reasoning based on Multi-Agent Large Language Model System
A schema-guided multi-agent LLM system called SG2 solves scene-graph Q&A and planning tasks by iteratively writing code to fetch only relevant graph data.
-
Psychometric-Based Evaluation for Theorem Proving with Large Language Models
The authors annotate miniF2F theorems with LLM-computed difficulty and discrimination scores, then use adaptive testing to rank 10 theorem-proving LLMs using only about 23% of the theorems.
-
Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions
LLM error detectors favor conventional solution formats; generating an adaptive reference solution before grading mitigates this conformity bias on a 200-example GSM8K subset.
-
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
DSTC builds preference pairs by selecting the hardest self-generated test that the best self-generated code passes, then fine-tunes code LMs with DPO or KTO to improve pass@1 accuracy.
-
Leveraging LLMs for Predictive Insights in Food Policy and Behavioral Interventions
A fine-tuned GPT-3.5 Turbo model predicts the direction of held-out food-policy experiments with 79% accuracy, but only 55% on preregistered unpublished studies.
-
Feature Generation Using LLMs: An Evolutionary Algorithm Approach
A funsearch-style evolutionary loop using LLaMA-3.1 7B-generated Python expressions creates new table features and improves F1 in 13 of 16 evaluated classification settings.
-
From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
Post-training LLMs first on abstract number-free reasoning plans (CoMT), then with confidence-weighted rewards (CCRL), raises math accuracy by ~2-5 points over standard CoT-SFT+RL across four models.
-
A Fragile Number Sense: Probing the Elemental Limits of Numerical Reasoning in LLMs
LLMs solve routine arithmetic and primality checks well but consistently fail at the Game of 24, which requires trial-and-error search.
-
A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis
An LLM agent that reframes beam analysis as OpenSeesPy code generation reaches over 99 percent reliability on a small benchmark, but chiefly because the prompt contains a near-identical solved example.
-
Advanced For-Loop for QML algorithm search
The paper sketches an LLM-based multi-agent framework that generated quantum variants of MLP, forward-forward, and backpropagation, but with no reproducible evidence that the search works.
-
WIP: Large Language Model-Enhanced Smart Tutor for Undergraduate Circuit Analysis
A GPT-4o-based smart tutor with problem-specific documents provided homework feedback in a circuit analysis course, and 90.9% of 66 student feedback responses were positive.
-
Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration
The paper proposes a systematic classification and two roadmaps for using foundation models (LLMs and wireless foundation models) to design Synesthesia of Machines systems for 6G, with preliminary case-study evidence ...
-
Understanding LLM Scientific Reasoning through Promptings and Model's Explanation on the Answers
On the GPQA benchmark, GPT-4o's highest accuracy came from self-consistency prompting, about 53 percent correct, but its explanations were least similar to the reference solutions, while direct answer and chain-of-tho...
-
Enhancing LLM Character-Level Manipulation via Divide and Conquer
ToCAD, a three-stage divide-and-conquer prompt, atomizes words into spaced letters, edits them, and reconstructs, sharply improving LLM exact-match accuracy on deletion, insertion, and substitution tasks.
-
Multiple Abstraction Level Retrieve Augment Generation
MAL-RAG retrieves document, section, paragraph, and multi-sentence chunks together and claims a 25.7% improvement in AI-judged answer correctness on glycoscience questions over single-level RAG.
-
LemmaHead: RAG Assisted Proof Generation Using Large Language Models
RAG with five rounds of iterative hint retrieval lets GPT-4 prove 40% of MiniF2F formal problems, up from a 9.4% single-pass baseline.
Discussion (0). Continue with ORCID to comment.