REVIEW 14 cited by
Impact of Pretraining Term Frequencies on Few-Shot Reasoning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Pretrained Language Models (LMs) have demonstrated ability to perform numerical reasoning by extrapolating from a few examples in few-shot settings. However, the extent to which this extrapolation relies on robust reasoning is unclear. In this paper, we investigate how well these models reason with terms that are less frequent in the pretraining data. In particular, we examine the correlations between the model performance on test instances and the frequency of terms from those instances in the pretraining data. We measure the strength of this correlation for a number of GPT-based language models (pretrained on the Pile dataset) on various numerical deduction tasks (e.g., arithmetic and unit conversion). Our results consistently demonstrate that models are more accurate on instances whose terms are more prevalent, in some cases above $70\%$ (absolute) more accurate on the top 10\% frequent terms in comparison to the bottom 10\%. Overall, although LMs exhibit strong performance at few-shot numerical reasoning tasks, our results raise the question of how much models actually generalize beyond pretraining data, and we encourage researchers to take the pretraining data into account when interpreting evaluation results.
Forward citations
Cited by 14 Pith papers
-
On the Fitness Landscape in the $NK$ Model
For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.
-
Strategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory
Frontier LLMs survive and often thrive in evolutionary Prisoner's Dilemma tournaments, and each model family shows a distinct, context-dependent cooperation fingerprint.
-
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
Training a language model by distilling a coach's written experiential knowledge beats training on a scalar rubric score for open-ended tasks, with better out-of-distribution transfer.
-
MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference
A new automatic method for creating minimal reasoning-preserving variants of NLI problems shows that 14 models drop 4 to 20 percent in accuracy on those variants.
-
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
An adaptive, per-language data filtering and deduplication pipeline produces multilingual LLM pre-training corpora that beat prior public datasets on 11 of 14 evaluated languages, and a 20TB, 1,868 language-script dat...
-
ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers
LLMs perform consistently better on tasks where input words are replaced with a consistent, reversible substitution cipher than when replacements are random, and the authors propose this gap as a measure of task learn...
-
UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models
UGMathBench provides a 5,062-problem dynamic benchmark for undergraduate math reasoning and shows that leading LLMs solve all versions of only about half the problems.
-
Randomly Sampled Language Reasoning Problems Elucidate Limitations of In-Context Learning
On randomly sampled 3-state DFA language tasks, foundation LLMs underperform n-gram baselines under pure in-context-learning prompts.
-
On the Reasoning Capacity of AI Models and How to Quantify It
Positional randomization on GPQA shows GPT-4o-mini's accuracy is inflated by position-dependent heuristics, but the paper's strategy-decomposition model is validated only by construction and contradicts its own accuracy data.
-
Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective
Task Decodability, a k-NN measure of how separable a task is in a model's middle-layer representations, tracks and predicts in-context learning accuracy, and early-layer finetuning improves it more than late-layer finetuning.
-
VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction
VASCAR uses GPT-4o and Gemini to iteratively refine poster layouts from rendered bounding-box images, achieving strong scores on PKU and CGL without training.
-
Understanding Chain-of-Thought in LLMs through Information Theory
A chain-of-thought step's information gain, estimated by a fine-tuned supervisor model, can identify the step where an LLM's reasoning first diverges from the correct answer.
-
Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models
DALL-E 3 fails to reliably produce images matching prompts about physical relations, negations, and exact numbers beyond three, and a grounded diffusion pipeline does worse on relations.
-
Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
A 4B-parameter model is claimed to explain its own reasoning through inverse attention analysis, but the paper offers no consistent evidence or artifacts.
Discussion (0). Continue with ORCID to comment.