Pith. sign in

REVIEW 14 cited by

Impact of Pretraining Term Frequencies on Few-Shot Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.07206 v2 pith:43PA5OQC submitted 2022-02-15 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelspretrainingdatareasoningtermsfew-shotinstancesnumerical
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Pretrained Language Models (LMs) have demonstrated ability to perform numerical reasoning by extrapolating from a few examples in few-shot settings. However, the extent to which this extrapolation relies on robust reasoning is unclear. In this paper, we investigate how well these models reason with terms that are less frequent in the pretraining data. In particular, we examine the correlations between the model performance on test instances and the frequency of terms from those instances in the pretraining data. We measure the strength of this correlation for a number of GPT-based language models (pretrained on the Pile dataset) on various numerical deduction tasks (e.g., arithmetic and unit conversion). Our results consistently demonstrate that models are more accurate on instances whose terms are more prevalent, in some cases above $70\%$ (absolute) more accurate on the top 10\% frequent terms in comparison to the bottom 10\%. Overall, although LMs exhibit strong performance at few-shot numerical reasoning tasks, our results raise the question of how much models actually generalize beyond pretraining data, and we encourage researchers to take the pretraining data into account when interpreting evaluation results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 52 citations worldwide. Full citation record

  1. On the Fitness Landscape in the $NK$ Model

    math.PR 2025-08 unverdicted novelty 7.0 of 10

    For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.

  2. Strategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory

    cs.AI 2025-07 conditional novelty 7.0 of 10

    Frontier LLMs survive and often thrive in evolutionary Prisoner's Dilemma tournaments, and each model family shows a distinct, context-dependent cooperation fingerprint.

  3. LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Training a language model by distilling a coach's written experiential knowledge beats training on a scalar rubric score for open-ended tasks, with better out-of-distribution transfer.

  4. MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference

    cs.CL 2025-10 conditional novelty 6.0 of 10

    A new automatic method for creating minimal reasoning-preserving variants of NLI problems shows that 14 models drop 4 to 20 percent in accuracy on those variants.

  5. FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An adaptive, per-language data filtering and deduplication pipeline produces multilingual LLM pre-training corpora that beat prior public datasets on 11 of 14 evaluated languages, and a 20TB, 1,868 language-script dat...

  6. ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers

    cs.CL 2025-04 conditional novelty 6.0 of 10

    LLMs perform consistently better on tasks where input words are replaced with a consistent, reversible substitution cipher than when replacements are random, and the authors propose this gap as a measure of task learn...

  7. UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    UGMathBench provides a 5,062-problem dynamic benchmark for undergraduate math reasoning and shows that leading LLMs solve all versions of only about half the problems.

  8. Randomly Sampled Language Reasoning Problems Elucidate Limitations of In-Context Learning

    cs.LG 2025-01 conditional novelty 6.0 of 10

    On randomly sampled 3-state DFA language tasks, foundation LLMs underperform n-gram baselines under pure in-context-learning prompts.

  9. On the Reasoning Capacity of AI Models and How to Quantify It

    cs.AI 2025-01 reject novelty 5.0 of 10

    Positional randomization on GPQA shows GPT-4o-mini's accuracy is inflated by position-dependent heuristics, but the paper's strategy-decomposition model is validated only by construction and contradicts its own accuracy data.

  10. Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Task Decodability, a k-NN measure of how separable a task is in a model's middle-layer representations, tracks and predicts in-context learning accuracy, and early-layer finetuning improves it more than late-layer finetuning.

  11. VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction

    cs.CV 2024-12 conditional novelty 5.0 of 10

    VASCAR uses GPT-4o and Gemini to iteratively refine poster layouts from rendered bounding-box images, achieving strong scores on PKU and CGL without training.

  12. Understanding Chain-of-Thought in LLMs through Information Theory

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A chain-of-thought step's information gain, estimated by a fine-tuned supervisor model, can identify the step where an LLM's reasoning first diverges from the correct answer.

  13. Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models

    cs.CV 2024-11 conditional novelty 4.0 of 10

    DALL-E 3 fails to reliably produce images matching prompts about physical relations, negations, and exact numbers beyond three, and a grounded diffusion pipeline does worse on relations.

  14. Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models

    cs.AI 2025-06 reject novelty 3.0 of 10

    A 4B-parameter model is claimed to explain its own reasoning through inverse attention analysis, but the paper offers no consistent evidence or artifacts.

Pith tools