Pith. sign in
Pith Number

pith:FJNRNXM5

pith:2022:FJNRNXM56FQ5232WSXQWFHHIY4
not attested not anchored not stored refs resolved

In-context Learning and Induction Heads

Amanda Askell, Andy Jones, Anna Chen, Ben Mann, Catherine Olsson, Chris Olah, Danny Hernandez, Dario Amodei, Dawn Drain, Deep Ganguli, Jack Clark, Jackson Kernion, Jared Kaplan, Kamal Ndousse, Liane Lovitt, Neel Nanda, Nelson Elhage, Nicholas Joseph, Nova DasSarma, Sam McCandlish, Scott Johnston, Tom Brown, Tom Conerly, Tom Henighan, Yuntao Bai, Zac Hatfield-Dodds

Induction heads implement the core copying algorithm behind in-context learning in transformers.

arxiv:2209.11895 v1 · 2022-09-24 · cs.LG

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{FJNRNXM56FQ5232WSXQWFHHIY4}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

induction heads might constitute the mechanism for the majority of all 'in-context learning' in large transformer models (i.e. decreasing loss at increasing token indices)

C2weakest assumption

That the emergence of induction heads is causally responsible for the observed increase in in-context learning ability rather than both phenomena being downstream effects of some other training dynamic.

C3one line summary

Induction heads, which implement pattern completion in attention, develop at the same training stage as a sudden rise in in-context learning, providing evidence they are the primary mechanism for in-context learning in transformers.

References

25 extracted · 25 resolved · 8 Pith anchors

[1] Language Models are Few-Shot Learners 2005 · arXiv:2005.14165
[2] Evaluating Large Language Models Trained on Code · arXiv:2107.03374
[3] arXiv preprint arXiv:2001.09977 , year= 2001
[4] Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets · arXiv:2201.02177
[5] Scaling Laws for Neural Language Models 2001 · arXiv:2001.08361

Cited by

160 papers in Pith

Receipt and verification
First computed 2026-07-05T05:00:29.389977Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

2a5b16dd9df161dd6f5695e1629ce8c733369779ac09705c36369356641a2981

Aliases

arxiv: 2209.11895 · arxiv_version: 2209.11895v1 · doi: 10.48550/arxiv.2209.11895 · pith_short_12: FJNRNXM56FQ5 · pith_short_16: FJNRNXM56FQ5232W · pith_short_8: FJNRNXM5
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/FJNRNXM56FQ5232WSXQWFHHIY4 \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 2a5b16dd9df161dd6f5695e1629ce8c733369779ac09705c36369356641a2981
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "a9d9d8560c4381d747a14e1cdf6e26154df31a0faafb25eba1a2760b94c23879",
    "cross_cats_sorted": [],
    "license": "http://creativecommons.org/licenses/by/4.0/",
    "primary_cat": "cs.LG",
    "submitted_at": "2022-09-24T00:43:19Z",
    "title_canon_sha256": "f8b2a0a341303093e25b2e1d2834fdfca6fa832a89c63e46757e7ceabc1cc78c"
  },
  "schema_version": "1.0",
  "source": {
    "id": "2209.11895",
    "kind": "arxiv",
    "version": 1
  }
}