Pith. sign in

REVIEW 8 cited by

Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.00871 v1 pith:4B6M3DR2 submitted 2023-11-01 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords pretrainingdatamodelsin-contexttaskscapabilitiesfamiliesmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Transformer models, notably large language models (LLMs), have the remarkable ability to perform in-context learning (ICL) -- to perform new tasks when prompted with unseen input-output examples without any explicit model training. In this work, we study how effectively transformers can bridge between their pretraining data mixture, comprised of multiple distinct task families, to identify and learn new tasks in-context which are both inside and outside the pretraining distribution. Building on previous work, we investigate this question in a controlled setting, where we study transformer models trained on sequences of $(x, f(x))$ pairs rather than natural language. Our empirical results show transformers demonstrate near-optimal unsupervised model selection capabilities, in their ability to first in-context identify different task families and in-context learn within them when the task families are well-represented in their pretraining data. However when presented with tasks or functions which are out-of-domain of their pretraining data, we demonstrate various failure modes of transformers and degradation of their generalization for even simple extrapolation tasks. Together our results highlight that the impressive ICL abilities of high-capacity sequence models may be more closely tied to the coverage of their pretraining data mixtures than inductive biases that create fundamental generalization capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Induction Heads Interpolate N-Grams

    cs.LG 2026-07 accept novelty 7.0 of 10

    Induction-head circuits implement soft context-matching (Jelinek–Mercer-style interpolation over partial matches) plus BOS-induced Dirichlet pseudo-counts, and trained transformers recover both mechanisms.

  2. How Context Attribution Handles What the Model Already Knows

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Context attribution methods cannot disentangle in-context from in-weight knowledge and assign unfaithful scores under overlap; new metrics and WMDP-Cyber++ quantify the failure.

  3. Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers

    cs.CL 2026-01 conditional novelty 6.0 of 10

    In a two-modality transformer, a primary-modality pretraining stage installs an induction circuit, so the secondary modality needs only low class diversity to learn in-context from examples.

  4. How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-Off

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Heavy-tailed pretraining distributions improve in-context task selection under distribution shift but worsen ICL generalization, especially in low-data regimes.

  5. Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention

    cs.CL 2025-09 conditional novelty 6.0 of 10

    ICR extracts shared attention directions from in-context learning and routes them at inference time, enabling zero-shot reuse across tasks.

  6. Selective Induction Heads: How Transformers Select Causal Structures In Context

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Transformers can learn to select the correct lag of an interleaved Markov chain in context via a circuit the authors call a selective induction head, whose asymptotic optimality proof is incomplete.

  7. Meta-Learning Approaches for Speaker-Dependent Voice Fatigue Models

    cs.LG 2025-05 reject novelty 6.0 of 10

    Meta-learning, especially a transformer-based sequence model, outperforms cross-sectional and mixed-effects baselines for predicting time since sleep from speech, though the evaluation protocol may overstate deploymen...

  8. Reverse Convolution and Its Applications to Image Restoration

    cs.CV 2025-08 reject novelty 4.0 of 10

    The abstract and body of this submission are two unrelated papers; the reverse-convolution claims appear nowhere in the full text.

Pith tools