Pith. sign in

REVIEW 10 cited by

A Theory of Emergent In-Context Learning as Implicit Structure Induction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07971 v1 pith:M7JED7IX submitted 2023-03-14 cs.CL cs.LG

classification cs.CLcs.LG
keywords in-contextlearningtheoreticalcompositionallanguageemergentllmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scaling large language models (LLMs) leads to an emergent capacity to learn in-context from example demonstrations. Despite progress, theoretical understanding of this phenomenon remains limited. We argue that in-context learning relies on recombination of compositional operations found in natural language data. We derive an information-theoretic bound showing how in-context learning abilities arise from generic next-token prediction when the pretraining distribution has sufficient amounts of compositional structure, under linguistically motivated assumptions. A second bound provides a theoretical justification for the empirical success of prompting LLMs to output intermediate steps towards an answer. To validate theoretical predictions, we introduce a controlled setup for inducing in-context learning; unlike previous approaches, it accounts for the compositional nature of language. Trained transformers can perform in-context learning for a range of tasks, in a manner consistent with the theoretical results. Mirroring real-world LLMs in a miniature setup, in-context learning emerges when scaling parameters and data, and models perform better when prompted to output intermediate steps. Probing shows that in-context learning is supported by a representation of the input's compositional structure. Taken together, these results provide a step towards theoretical understanding of emergent behavior in large language models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Agentic Transformers Provably Learn to Search via Reinforcement Learning

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    In a stochastic k-ary tree, a two-head transformer learns randomized DFS via policy gradient under depth-wise curriculum, generalizes to deeper trees, and adapts to imbalanced goals via discounting.

  2. Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Manifold steering along activation geometry induces behavioral trajectories matching the natural manifold of outputs, while linear steering produces off-manifold unnatural behaviors.

  3. Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Hawk raises NPU kernel generation accuracy from 49.4% to 80% and yields up to 2.2× speedups by retrieving and distilling structured hardware-aware knowledge without any model training.

  4. Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

    cs.AI 2026-07 unverdicted novelty 6.0 of 10

    Hawk is a training-free framework that boosts NPU kernel generation accuracy to 80% and achieves up to 2.2x speedup via hardware-aware knowledge synthesis, 2D retrieval, and effect-driven distillation.

  5. Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    Humans and LLMs exhibit similar error patterns in common-sense reasoning, consistent with shared pattern-matching mechanisms rather than abstract world models.

  6. Learning to Remember, Learn, and Forget in Attention-Based Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Palimpsa adds a per-slot importance/precision state to gated linear attention, letting a fixed-size memory forget stale information and protect important information, and recovers Mamba2 as a high-forgetting limit.

  7. Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models

    cs.CL 2026-02 unverdicted novelty 6.0 of 10

    LLMs dynamically construct and causally rely on structured conceptual subspaces in middle-to-late layers for in-context inference.

  8. Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training

    cs.LG 2025-11 conditional novelty 6.0 of 10

    Curriculum post-training on reasoning trees yields polynomial sample complexity for accurate Chain-of-Thought generation in transformers, unlike exponential requirements without curriculum.

  9. InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A submodular mutual information framework for selecting and training in-context learning exemplars improves average accuracy on nine benchmarks by about five points over the IDEAL baseline.

  10. A Survey on In-context Learning

    cs.CL 2022-12 unverdicted novelty 3.0 of 10

    The paper surveys definitions, techniques, applications, and challenges in in-context learning for large language models.

Pith tools