Pith. sign in

REVIEW 11 cited by

How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.00586 v5 pith:SRPBUEJT submitted 2023-04-30 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords gpt-2smallabilitiescircuitlanguagemathematicalpre-trainedyear
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Pre-trained language models can be surprisingly adept at tasks they were not explicitly trained on, but how they implement these capabilities is poorly understood. In this paper, we investigate the basic mathematical abilities often acquired by pre-trained language models. Concretely, we use mechanistic interpretability techniques to explain the (limited) mathematical abilities of GPT-2 small. As a case study, we examine its ability to take in sentences such as "The war lasted from the year 1732 to the year 17", and predict valid two-digit end years (years > 32). We first identify a circuit, a small subset of GPT-2 small's computational graph that computes this task's output. Then, we explain the role of each circuit component, showing that GPT-2 small's final multi-layer perceptrons boost the probability of end years greater than the start year. Finally, we find related tasks that activate our circuit. Our results suggest that GPT-2 small computes greater-than using a complex but general mechanism that activates across diverse contexts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Alignment on Llama 3 reduces explicit bias but amplifies implicit bias, because aligned models no longer represent 'black' and 'white' as racial concepts in ambiguous contexts.

  2. Interpreting Language Model Hidden States at Scale

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A single low-rank, sparsely trained lens family decodes residual, attention, and MLP states in models up to 70B parameters, revealing that visible and causally effective locations for a behavior can differ.

  3. Pretraining Curricula Enable Selective Fine-tuning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Imbalanced pretraining curricula disentangle task circuits in transformers, improving in-context learning and the selectivity of refusal fine-tuning relative to balanced training.

  4. Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning

    cs.LG 2026-01 reject novelty 6.0 of 10

    Spectral features of attention are claimed to classify proof validity with near-perfect effect sizes, but the main evaluation relabels proofs using the classifier's own outputs.

  5. On Mechanistic Circuits for Extractive Question-Answering

    cs.CL 2025-02 conditional novelty 6.0 of 10

    One attention head from the extracted context-faithfulness circuit provides reliable extractive QA attribution and improves context faithfulness when its attributions are added to the prompt.

  6. Structure Development in List-Sorting Transformers

    cs.LG 2025-01 conditional novelty 6.0 of 10

    In a list-sorting transformer, the mean and variance of gaps between adjacent sorted numbers determine whether attention heads split the vocabulary, suppress copying to calibrate, or switch off.

  7. ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis

    cs.CL 2024-12 conditional novelty 6.0 of 10

    ResoFilter keeps fine-tuning examples that produce small parameter updates in the last layers, matching full-data fine-tuning on GSM8k with 50% of the math data.

  8. Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability

    cs.LG 2024-11 conditional novelty 6.0 of 10

    GPT-2 small's IOI circuit adapts to new prompt formats mostly by reusing components, but the base circuit's apparent success is partly an evaluation artifact called S2 Hacking.

  9. From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits

    cs.CL 2025-08 conditional novelty 5.0 of 10

    GPT-2 small performs syllogisms through truth-copying attention heads and a suppression-plus-MLP pathway that can output a negated truth value.

  10. NEAT: Concept driven Neuron Attribution in LLMs

    cs.CL 2025-08 reject novelty 4.0 of 10

    NEAT identifies concept neurons by feeding a single mean hidden-state vector through the model and ranking neurons by their effect on concept-word probabilities.

  11. Towards a Theory of AI Personhood

    cs.AI 2025-01 accept novelty 4.0 of 10

    The paper outlines agency, theory of mind, and self-awareness as necessary conditions for AI personhood, reviews inconclusive evidence, and argues that AI personhood would make control-focused alignment ethically problematic.

Pith tools