Pith. sign in

REVIEW 11 cited by

A Theory of Emergent In-Context Learning as Implicit Structure Induction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07971 v1 pith:M7JED7IX submitted 2023-03-14 cs.CL cs.LG

classification cs.CLcs.LG
keywords in-contextlearningtheoreticalcompositionallanguageemergentllmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Scaling large language models (LLMs) leads to an emergent capacity to learn in-context from example demonstrations. Despite progress, theoretical understanding of this phenomenon remains limited. We argue that in-context learning relies on recombination of compositional operations found in natural language data. We derive an information-theoretic bound showing how in-context learning abilities arise from generic next-token prediction when the pretraining distribution has sufficient amounts of compositional structure, under linguistically motivated assumptions. A second bound provides a theoretical justification for the empirical success of prompting LLMs to output intermediate steps towards an answer. To validate theoretical predictions, we introduce a controlled setup for inducing in-context learning; unlike previous approaches, it accounts for the compositional nature of language. Trained transformers can perform in-context learning for a range of tasks, in a manner consistent with the theoretical results. Mirroring real-world LLMs in a miniature setup, in-context learning emerges when scaling parameters and data, and models perform better when prompted to output intermediate steps. Probing shows that in-context learning is supported by a representation of the input's compositional structure. Taken together, these results provide a step towards theoretical understanding of emergent behavior in large language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars

    cs.CL 2025-04 conditional novelty 7.0 of 10

    Prepending metadata during pre-training helps language models when downstream prompts are long enough to infer the underlying semantics, but hurts when prompts are short.

  2. Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

    cs.AI 2026-07 unverdicted novelty 6.0 of 10

    Hawk raises NPU kernel generation accuracy from 49.4% to 80% and yields up to 2.2× speedups by retrieving and distilling structured hardware-aware knowledge without any model training.

  3. Learning to Remember, Learn, and Forget in Attention-Based Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Palimpsa adds a per-slot importance/precision state to gated linear attention, letting a fixed-size memory forget stale information and protect important information, and recovers Mamba2 as a high-forgetting limit.

  4. The Other Mind: How Language Models Exhibit Human Temporal Cognition

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.

  5. Next-Token Prediction Should be Ambiguity-Sensitive: A Meta-Learning Perspective

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Transformers systematically deviate from the Bayes-optimal predictor under high-ambiguity contexts on a new HMM benchmark, and a Monte Carlo predictor that decouples task inference from token prediction partly closes ...

  6. Transformers Meet In-Context Learning: A Universal Approximation Theory

    cs.LG 2025-06 accept novelty 6.0 of 10

    A constructive theorem shows that transformers can perform in-context learning for any Barron-type function class by combining universal features with an emulated Lasso solver.

  7. ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers

    cs.CL 2025-04 conditional novelty 6.0 of 10

    LLMs perform consistently better on tasks where input words are replaced with a consistent, reversible substitution cipher than when replacements are random, and the authors propose this gap as a measure of task learn...

  8. IC-Cache: Efficient Large Language Model Serving via In-context Caching

    cs.LG 2025-01 conditional novelty 6.0 of 10

    IC-Cache reuses historical large-model responses as in-context examples so small models can handle a larger share of serving traffic without losing quality, improving throughput and latency.

  9. InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A submodular mutual information framework for selecting and training in-context learning exemplars improves average accuracy on nine benchmarks by about five points over the IDEAL baseline.

  10. Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning

    cs.LG 2025-06 reject novelty 3.0 of 10

    In-context learning is reframed as implicit knowledge distillation, but the main claims either restate the known attention-equals-gradient-descent result or build the prompt-shift bound into the definition of MMD.

  11. A Survey on Large Language Models with some Insights on their Capabilities and Limitations

    cs.CL 2025-01 unverdicted novelty 3.0 of 10

    A broad survey of LLM methods and applications, plus an empirical section on how code-rich pretraining may influence chain-of-thought reasoning, the details of which are not visible in the supplied text.

Pith tools