Pith. sign in

REVIEW 2 cited by

Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.00053 v3 pith:4STLGXCX submitted 2024-05-28 cs.CL cs.LG

classification cs.CLcs.LG
keywords in-contextlearningmodelmodelsstructurallanguageabilityin-weights
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Language models have the ability to perform in-context learning (ICL), allowing them to flexibly adapt their behavior based on context. This contrasts with in-weights learning (IWL), where memorized information is encoded in model parameters after iterated observations of data. An ideal model should be able to flexibly deploy both of these abilities. Despite their apparent ability to learn in-context, language models are known to struggle when faced with unseen or rarely seen tokens (Land & Bartolo, 2024). Hence, we study $\textbf{structural in-context learning}$, which we define as the ability of a model to execute in-context learning on arbitrary novel tokens -- so called because the model must generalize on the basis of e.g. sentence structure or task structure, rather than content encoded in token embeddings. We study structural in-context algorithms on both synthetic and naturalistic tasks using toy models, masked language models, and autoregressive language models. We find that structural ICL appears before quickly disappearing early in LM pretraining. While it has been shown that ICL can diminish during training (Singh et al., 2023), we find that prior work does not account for structural ICL. Building on Chen et al. (2024) 's active forgetting method, we introduce pretraining and finetuning methods that can modulate the preference for structural ICL and IWL. Importantly, this allows us to induce a $\textit{dual process strategy}$ where in-context and in-weights solutions coexist within a single model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In-Context Learning Strategies Emerge Rationally

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Transformer in-context learning is modeled as a posterior-weighted mixture of memorizing and generalizing Bayesian predictors, with a loss-complexity tradeoff governed by three fitted parameters.

  2. ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers

    cs.CL 2025-04 conditional novelty 6.0 of 10

    LLMs perform consistently better on tasks where input words are replaced with a consistent, reversible substitution cipher than when replacements are random, and the authors propose this gap as a measure of task learn...

Pith tools