{"paper":{"title":"In-context Learning and Induction Heads","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Induction heads implement the core copying algorithm behind in-context learning in transformers.","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Amanda Askell, Andy Jones, Anna Chen, Ben Mann, Catherine Olsson, Chris Olah, Danny Hernandez, Dario Amodei, Dawn Drain, Deep Ganguli, Jack Clark, Jackson Kernion, Jared Kaplan, Kamal Ndousse, Liane Lovitt, Neel Nanda, Nelson Elhage, Nicholas Joseph, Nova DasSarma, Sam McCandlish, Scott Johnston, Tom Brown, Tom Conerly, Tom Henighan, Yuntao Bai, Zac Hatfield-Dodds","submitted_at":"2022-09-24T00:43:19Z","abstract_excerpt":"\"Induction heads\" are attention heads that implement a simple algorithm to complete token sequences like [A][B] ... [A] -> [B]. In this work, we present preliminary and indirect evidence for a hypothesis that induction heads might constitute the mechanism for the majority of all \"in-context learning\" in large transformer models (i.e. decreasing loss at increasing token indices). We find that induction heads develop at precisely the same point as a sudden sharp increase in in-context learning ability, visible as a bump in the training loss. We present six complementary lines of evidence, arguin"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"induction heads might constitute the mechanism for the majority of all 'in-context learning' in large transformer models (i.e. decreasing loss at increasing token indices)","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the emergence of induction heads is causally responsible for the observed increase in in-context learning ability rather than both phenomena being downstream effects of some other training dynamic.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Induction heads, which implement pattern completion in attention, develop at the same training stage as a sudden rise in in-context learning, providing evidence they are the primary mechanism for in-context learning in transformers.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Induction heads implement the core copying algorithm behind in-context learning in transformers.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"82e3dd5268fe1c82b5db72e78fda60ad9d8712d0d81a9f7134c3f9c55571f66e"},"source":{"id":"2209.11895","kind":"arxiv","version":1},"verdict":{"id":"b3aad1ed-d955-4cb3-a3c7-d89ce6c0dd91","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-11T03:44:13.815348Z","strongest_claim":"induction heads might constitute the mechanism for the majority of all 'in-context learning' in large transformer models (i.e. decreasing loss at increasing token indices)","one_line_summary":"Induction heads, which implement pattern completion in attention, develop at the same training stage as a sudden rise in in-context learning, providing evidence they are the primary mechanism for in-context learning in transformers.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the emergence of induction heads is causally responsible for the observed increase in in-context learning ability rather than both phenomena being downstream effects of some other training dynamic.","pith_extraction_headline":"Induction heads implement the core copying algorithm behind in-context learning in transformers."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2209.11895/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":25,"sample":[{"doi":"","year":2005,"title":"Language Models are Few-Shot Learners","work_id":"214732c0-2edd-44a0-af9e-28184a2b8279","ref_index":1,"cited_arxiv_id":"2005.14165","is_internal_anchor":true},{"doi":"","year":null,"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","ref_index":2,"cited_arxiv_id":"2107.03374","is_internal_anchor":true},{"doi":"","year":2001,"title":"arXiv preprint arXiv:2001.09977 , year=","work_id":"ba5130eb-69cb-4786-944d-f724b5f7876c","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":null,"title":"Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets","work_id":"a3c30ead-1625-4c18-a9c1-e4928dcd0da6","ref_index":4,"cited_arxiv_id":"2201.02177","is_internal_anchor":true},{"doi":"","year":2001,"title":"Scaling Laws for Neural Language Models","work_id":"b7dd8749-9c45-4977-ab9b-64478dce1ae8","ref_index":5,"cited_arxiv_id":"2001.08361","is_internal_anchor":true}],"resolved_work":25,"snapshot_sha256":"0a95b95e3169f058af2f8400eb5b907b26253ed25e786e056a4b4e1e051862db","internal_anchors":8},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}