REVIEW 17 cited by
Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advances in language models have demonstrated their capability to solve mathematical reasoning problems, achieving near-perfect accuracy on grade-school level math benchmarks like GSM8K. In this paper, we formally study how language models solve these problems. We design a series of controlled experiments to address several fundamental questions: (1) Can language models truly develop reasoning skills, or do they simply memorize templates? (2) What is the model's hidden (mental) reasoning process? (3) Do models solve math questions using skills similar to or different from humans? (4) Do models trained on GSM8K-like datasets develop reasoning skills beyond those necessary for solving GSM8K problems? (5) What mental process causes models to make reasoning mistakes? (6) How large or deep must a model be to effectively solve GSM8K-level math questions? Our study uncovers many hidden mechanisms by which language models solve mathematical questions, providing insights that extend beyond current understandings of LLMs.
Forward citations
Cited by 17 Pith papers
-
Theoretical limitations of multi-layer Transformer
An L-layer decoder-only Transformer requires polynomial model dimension to compute L-step sequential function composition, and this is proven without any unproven complexity conjecture.
-
Understanding Reasoning from Pretraining to Post-Training
A joint scaling law: post-RL chess and math performance is predictable from pretraining loss, RL improvement rate grows with pretraining tokens, and RL both amplifies and discovers moves depending on difficulty.
-
How Data Shapes RoPE Frequency Usage: From Positional Scale Matching to Length Generalization
RoPE frequency usage is determined by a data-induced dependency width W, with the optimal frequency scaling as π/W, explaining both learned spectra and the success of position interpolation.
-
Scaling Latent Reasoning via Looped Language Models
Looped language models with latent iterative computation and entropy-regularized depth allocation achieve performance matching up to 12B standard LLMs through superior knowledge manipulation.
-
Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs
RL-trained LLMs keep most of their skills after weight merging, while SFT-trained LLMs drop about 19% on average, because RL keeps parameter updates smaller and more task-compatible.
-
The Power of Power Law: Asymmetry Enables Compositional Reasoning
Power-law data sampling creates beneficial asymmetry in the loss landscape that lets models acquire high-frequency skill compositions first, enabling more efficient learning of rare long-tail skills than uniform distr...
-
Depth Gives a False Sense of Privacy: LLM Internal States Inversion
LLM internal states at intermediate layers contain enough information to recover long, sensitive user prompts with high accuracy.
-
When More is Less: Understanding Chain-of-Thought Length in LLMs
LLM accuracy follows an inverted U in chain-of-thought length, with an optimal length that grows with task difficulty and shrinks with model capability.
-
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
A new synthetic benchmark reveals that LLM reasoning accuracy decays sigmoidally with problem complexity and that repeated sampling has poor scaling efficiency.
-
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel
StagFormer staggers transformer layers one time step apart with delayed cross-attention, enabling depth-parallel decoding at quality comparable to a deeper baseline.
-
Paradigm-Based Automatic HDL Code Generation Using LLMs
A paradigm-based workflow with information-list reuse and a two-phase loop improves LLM-generated Verilog pass rates on VerilogEval, with the full-dataset result built from a hybrid of baseline and proposed-method outputs.
-
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
ArxivRoll builds one-time private benchmark questions from recent arXiv papers and computes a rugged score that it claims estimates contamination and training bias in public LLM benchmarks.
-
Position: We Need An Algorithmic Understanding of Generative AI
The paper argues for a systematic algorithmic understanding of LLMs and presents a case study suggesting that Llama models do not implement BFS or DFS on graph navigation tasks.
-
Lightweight Latent Verifiers for Efficient Meta-Generation Strategies
LiLaVe, an XGBoost verifier trained on hidden states of a base LLM, predicts answer correctness with AUC comparable to large LLM-based verifiers, and enables conditional majority voting and self-correction that improv...
-
Spectral Journey: How Transformers Predict the Shortest Path
Two-layer transformers learn shortest paths on small graphs by building embeddings that correlate with spectral decomposition of the line graph, yielding an approximate spectral path-finding algorithm.
-
Analyzing Memorization in Large Language Models through the Lens of Model Attribution
Bypassing attention in the deepest transformer layers reduces extractable memorization with little loss on standard benchmarks, while early-layer bypass collapses the model.
-
Do LLMs Really Think Step-by-step In Implicit Reasoning?
A prompted LLM showed little linear-probing evidence of computing intermediate arithmetic steps during silent reasoning, whereas a model trained to internalize chain-of-thought did, and both degraded sharply under for...
Discussion (0). Continue with ORCID to comment.