Pith. sign in

REVIEW 3 major objections 2 minor 1 cited by

Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts

T0 review · 3 major / 2 minor · reviewed 2026-06-30 · grok-4.3

Pith's one-line read A sparse set of keystone neurons, isolated by consistent high activation across tasks, controls core capabilities in open-weight Transformers.

desk verdict The paper isolates a sparse cross-task activation subset whose removal collapses behavior and whose selective fine-tuning matches full updates, but the causal link to unique criticality still needs matched controls. read the letter →

arxiv 2605.24846 v2 pith:5OYWEVLQ submitted 2026-05-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords keystoneneuronsLLMinterpretabilityneuronactivationsparsesubsetsfine-tuningtransformermodelspretrainingcapabilitypreservation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper examines neuron activations inside large language models during inference on varied tasks. It identifies an extremely small subset that stays highly active no matter the task and shows that deleting this subset destroys model performance. These keystone neurons form early in pretraining and stay fixed, with their exact parameter values proving essential. The authors then demonstrate that fine-tuning only this subset produces gains equal to or better than updating every parameter while harming unrelated abilities less. A sympathetic reader would see this as evidence that models rely on a tiny, stable core rather than distributed computation across all neurons.

What carries the argument

Keystone neurons: the sparse subset isolated by high cross-task activation strength whose removal collapses model behavior.

What would settle it

An experiment in which randomly chosen neurons matched for activation strength produce the same collapse upon removal, or in which keystone neurons selected from one prompt set fail to affect held-out tasks.

Watch

Extended reading notes

Core claim

Across a wide range of open-weight Transformers, a subset of neurons remains consistently highly activated during inference across tasks of multiple capability dimensions. By probing along the cross-task activation strength, an extremely sparse subset is isolated, whose removal causes a collapse in model behavior, which we term keystone neurons. Our analysis reveals that keystone neurons are a stable and intrinsic neuron subset of the model that is largely established during pretraining. The parameters associated with these neurons are tightly calibrated during the training process, and their precise values are critical for the capabilities of the model.

Load-bearing premise

The high cross-task activation and the performance collapse after removal are caused by these neurons being uniquely critical rather than by the selection method itself or by other components that were not isolated.

Editorial extensions

If this is right

  • Updating only keystone neurons during supervised fine-tuning yields task gains comparable to or better than full-parameter fine-tuning.
  • Targeted updates on keystone neurons better preserve performance in other capability dimensions.
  • Keystone neurons form a stable intrinsic subset largely established during pretraining.
  • Precise parameter values tied to these neurons are critical for overall model capabilities.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The finding could enable prompt-only auditing of model internals without full retraining.
  • It raises the possibility that similar sparse critical subsets exist in non-Transformer architectures.
  • Targeted editing of keystone neurons might support more precise capability addition or removal.
  • The approach could extend to measuring how pretraining data distributions shape these stable subsets.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The manuscript claims that across open-weight Transformers, a sparse subset of neurons exhibits consistently high activation across tasks spanning multiple capability dimensions. By selecting on cross-task activation strength, an extremely sparse 'keystone' subset is isolated whose removal produces behavioral collapse; these neurons are argued to be stable, intrinsic, and largely fixed during pretraining with tightly calibrated parameters. The work further proposes a supervised fine-tuning procedure that updates only the keystone neurons and reports task gains comparable or superior to full-parameter fine-tuning while better preserving performance on unrelated capabilities despite modifying far fewer parameters.

Significance. If the empirical claims are substantiated with appropriate controls and quantitative results, the identification of a stable, pretraining-established sparse subset whose targeted update yields efficient adaptation would be a notable contribution to LLM interpretability and parameter-efficient fine-tuning. The potential reduction in modified parameters while maintaining or improving multi-task performance could influence both mechanistic understanding and practical deployment.

major comments (3)
  1. [Abstract / probing procedure] Abstract and method description: the central claim that ablation of the cross-task high-activation subset produces collapse specifically because of the cross-task consistency property (rather than generic high-activation or magnitude properties) requires explicit controls. No description is given of matched-cardinality random subsets, single-task activation subsets, or gradient-based importance baselines that would isolate the selection criterion; without these, the observed drop cannot be attributed to the stated mechanism.
  2. [Abstract / fine-tuning section] Abstract and experiments: the fine-tuning claim that updating only keystone neurons yields 'comparable or even better' task gains while better preserving other capabilities is presented without any reported metrics, number of updated parameters, task suite, baseline comparisons, or ablation on the selection threshold. These quantitative details are load-bearing for the practical contribution and are absent from the provided text.
  3. [Abstract / analysis of pretraining stability] Abstract: the assertion that keystone neurons are 'largely established during pretraining' and that their 'precise values are critical' is stated without supporting evidence such as activation statistics across training checkpoints, parameter-sensitivity analysis, or comparison to randomly initialized models. This is central to the intrinsic-property claim yet unsupported in the given material.
minor comments (2)
  1. [Method] Notation for activation strength and the precise definition of the 'probing along the cross-task activation strength' procedure should be formalized with an equation or algorithm box to allow replication.
  2. [Introduction] The term 'keystone neurons' is introduced without reference to prior related concepts in the interpretability literature (e.g., 'critical neurons' or 'superposition' studies); a brief related-work paragraph would clarify novelty.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for their valuable feedback on our manuscript. We address each of the major comments point by point below, and we will make the necessary revisions to strengthen the empirical support for our claims.

read point-by-point responses
  1. Referee: [Abstract / probing procedure] Abstract and method description: the central claim that ablation of the cross-task high-activation subset produces collapse specifically because of the cross-task consistency property (rather than generic high-activation or magnitude properties) requires explicit controls. No description is given of matched-cardinality random subsets, single-task activation subsets, or gradient-based importance baselines that would isolate the selection criterion; without these, the observed drop cannot be attributed to the stated mechanism.

    Authors: We agree that the manuscript requires these controls to properly attribute the effect to the cross-task consistency. The current version does not provide descriptions of matched-cardinality random subsets, single-task activation subsets, or gradient-based baselines. We will add these controls to the probing procedure and results in the revised manuscript. revision: yes

  2. Referee: [Abstract / fine-tuning section] Abstract and experiments: the fine-tuning claim that updating only keystone neurons yields 'comparable or even better' task gains while better preserving other capabilities is presented without any reported metrics, number of updated parameters, task suite, baseline comparisons, or ablation on the selection threshold. These quantitative details are load-bearing for the practical contribution and are absent from the provided text.

    Authors: The referee is correct that the provided manuscript text does not include the specific quantitative metrics, parameter counts, task suite details, baseline comparisons, or threshold ablations for the fine-tuning experiments. We will revise the experiments section to report these details comprehensively, including a table with the metrics and ablations. revision: yes

  3. Referee: [Abstract / analysis of pretraining stability] Abstract: the assertion that keystone neurons are 'largely established during pretraining' and that their 'precise values are critical' is stated without supporting evidence such as activation statistics across training checkpoints, parameter-sensitivity analysis, or comparison to randomly initialized models. This is central to the intrinsic-property claim yet unsupported in the given material.

    Authors: We acknowledge that the current manuscript does not include the supporting evidence such as activation statistics across checkpoints, parameter-sensitivity analysis, or comparisons to random initialization. We will perform and incorporate these analyses into the revised version to support the claims about pretraining stability and parameter criticality. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical selection and ablation claims remain independent of inputs

full rationale

The paper contains no equations, derivations, fitted parameters presented as predictions, or self-citation chains that reduce claims to their own inputs by construction. Keystone neuron identification proceeds from cross-task activation measurements followed by ablation experiments; these are observable, falsifiable steps whose outcomes are not forced by the selection criterion itself. The fine-tuning proposal similarly updates a pre-identified subset without renaming or smuggling prior results. No load-bearing step matches any enumerated circularity pattern.

Assumptions & free parameters 0 free parameters · 0 assumptions · 1 invented entities

Only abstract available; no explicit free parameters, axioms, or invented entities beyond the term 'keystone neurons' are stated.

invented entities (1)
  • keystone neurons
    purpose: Label for the sparse, cross-task highly activated subset whose removal collapses behavior
    New term introduced to describe the identified neurons; no independent evidence outside the paper's claimed experiments is mentioned.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts." pith.science (2026). https://pith.science/paper/5OYWEVLQ

@misc{pith2026260524846,
  author       = {Pith},
  title        = {Pith review of: Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5OYWEVLQ}},
  note         = {Machine review of arXiv:2605.24846}
}
read the original abstract

Large language models (LLMs) display strong comprehensive abilities, yet the internal mechanisms that support these behaviors remain insufficiently understood. In this work, we show that across a wide range of open-weight Transformers, a subset of neurons remains consistently highly activated during inference across tasks of multiple capability dimensions. By probing along the cross-task activation strength, an extremely sparse subset is isolated, whose removal causes a collapse in model behavior, which we term keystone neurons. Our analysis reveals that keystone neurons are a stable and intrinsic neuron subset of the model that is largely established during pretraining. The parameters associated with these neurons are tightly calibrated during the training process, and their precise values are critical for the capabilities of the model. Building on these insights, we propose a supervised fine-tuning approach that updates only keystone neurons, achieving task gains comparable to or even better than full-parameter fine-tuning while better preserving performance in other capability dimensions, despite modifying a much smaller number of parameters.

Figures

Figures reproduced from arXiv: 2605.24846 by the authors.

Figure 1
Figure 1. Conceptual illustration of keystone neurons. These neurons are consistently engaged across diverse prompts, and their deactivation leads to global capability collapse. rapid progress on external capabilities, the understanding of the internal mechanisms that give rise to these behaviors remains limited. A growing body of work shows that LLMs do not rely on all parameters in the same way, but instead develop stable p… view at source ↗
Figure 2
Figure 2. Neuron intersections and accuracy for Qwen2.5-7B-Instruct when selecting the top-α fraction of neurons ranked by activation across four capability dimensions. Bars show, for each α, how many neurons appear in the top-α sets for 1, 2, 3, or all 4 dimensions, while the black curve (right axis) reports the model comprehensive capability after masking the neurons in the 4-dimension intersection. model families, suggesti… view at source ↗
Figure 3
Figure 3. Performance under multiplicative rescaling of keystone vs. random neurons in Qwen2.5-7B-Instruct and Qwen2.5-0.5B-Instruct, measured by a fixed-weight aggregate score over MMLU, Math500, MGSM, and EvalPlus [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Mean–CV activation patterns of keystone neurons in different submodules of Llama-3.1-8B-Instruct. Each panel plots all neurons in one module (self-attention Q/K/V /O projections and FFN up/down projections) as blue dots in the plane of mean activation (x-axis) versus c…
Figure 5
Figure 5. Figure 5: Layer-wise counts of keystone neurons for base vs. instruction-tuned models across Qwen2.5, Qwen3-MoE, Llama3, and Gemma3 families; panel titles report base–instruct keystone overlap (%) [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

    cs.AI 2026-08 conditional novelty 7.0 of 10

    AgentOPSD converts sparse outcome rewards into turn-level credit by recursively accumulating self-distillation evidence in log-odds space, beating GRPO on most agentic benchmarks.

Reference graph

Works this paper leans on

53 extracted references · 53 canonical work pages · cited by 1 Pith paper

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...

  3. [3]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradb...

  4. [4]

    Gpt-4 technical report, 2024

    OpenAI. Gpt-4 technical report, 2024

  5. [5]

    Sparks of artificial general intelligence: Early experiments with gpt-4, 2023

    Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments with gpt-4, 2023

  6. [6]

    Solving quantitative reasoning problems with language models, 2022

    Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. Solving quantitative reasoning problems with language models, 2022

  7. [7]

    Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals

    Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushm...

  8. [8]

    A survey of large language models, 2025

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. A survey of large language models, 2025

Show all 53 references
  1. [9]

    Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

    DeepSeek - AI. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. page 633–638, September 2025

  2. [10]

    Longcat-flash-thinking-2601 technical report, 2026

    Meituan LongCat Team. Longcat-flash-thinking-2601 technical report, 2026

  3. [11]

    On generative agents in recommendation, 2024

    An Zhang, Yuxin Chen, Leheng Sheng, Xiang Wang, and Tat-Seng Chua. On generative agents in recommendation, 2024

  4. [12]

    Internalizing safety understanding in large reasoning models via verification, 2026

    Yi Zhang, Yuxin Chen, Leheng Sheng, Dongcheng Zhang, Chaochao Lu, Xiang Wang, and An Zhang. Internalizing safety understanding in large reasoning models via verification, 2026

  5. [13]

    L-mtp: Leap multi-token prediction beyond adjacent context for large language models, 2025

    Xiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang, Xianzhi Yu, Xiu Su, Shuo Yang, See-Kiong Ng, and Tat-Seng Chua. L-mtp: Leap multi-token prediction beyond adjacent context for large language models, 2025

  6. [14]

    Unlocking emergent modularity in large language models, 2024

    Zihan Qiu, Zeyu Huang, and Jie Fu. Unlocking emergent modularity in large language models, 2024

  7. [15]

    Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022

    William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022

  8. [16]

    St-moe: Designing stable and transferable sparse expert models, 2022

    Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. St-moe: Designing stable and transferable sparse expert models, 2022

  9. [17]

    Towards monosemanticity: Decomposing language models with dictionary learning, 2023

    Tom Bricken, Nelson Elhage, Chris Olah, Catherine Olsson, Nicholas Schiefer, Ben Mann, and Dario Amodei. Towards monosemanticity: Decomposing language models with dictionary learning, 2023. Transformer Circuits Thread

  10. [18]

    Llm modularity: The separability of capabilities in large language models, 2023

    Nicky Pochinkov. Llm modularity: The separability of capabilities in large language models, 2023. LessWrong / AI Alignment Forum

  11. [19]

    Language-specific neurons: The key to multilingual capabilities in large language models, 2024

    Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen. Language-specific neurons: The key to multilingual capabilities in large language models, 2024

  12. [20]

    The emergence of abstract thought in large language models beyond any language, 2025

    Yuxin Chen, Yiran Zhao, Yang Zhang, An Zhang, Kenji Kawaguchi, Shafiq Joty, Junnan Li, Tat-Seng Chua, Michael Qizhe Shieh, and Wenxuan Zhang. The emergence of abstract thought in large language models beyond any language, 2025

  13. [21]

    How do large language models handle multilingualism?, 2024

    Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, and Lidong Bing. How do large language models handle multilingualism?, 2024

  14. [22]

    Mechanistic understanding of language models in syntactic code completion, 2025

    Samuel Miller, Daking Rai, and Ziyu Yao. Mechanistic understanding of language models in syntactic code completion, 2025

  15. [23]

    Interpreting arithmetic mechanism in large language models through comparative neuron analysis, 2024

    Zeping Yu and Sophia Ananiadou. Interpreting arithmetic mechanism in large language models through comparative neuron analysis, 2024

  16. [24]

    Ran Song, Shizhu He, Shuting Jiang, Yantuan Xian, Shengxiang Gao, Kang Liu, and Zhengtao Yu. Does large language model contain task-specific neurons? In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natur...

  17. [25]

    Not all parameters are created equal: Smart isolation boosts fine-tuning performance, 2025 a

    Yao Wang, Di Liang, and Minlong Peng. Not all parameters are created equal: Smart isolation boosts fine-tuning performance, 2025 a

  18. [26]

    The lottery ticket hypothesis: Finding sparse, trainable neural networks, 2019

    Jonathan Frankle and Michael Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks, 2019

  19. [27]

    Investigating pattern neurons in urban time series forecasting

    Chengxin Wang, Yiran Zhao, Shaofeng Cai, and Gary Tan. Investigating pattern neurons in urban time series forecasting. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.net, 2025 b

  20. [28]

    Qwen3 technical report, 2025

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 40 others. Qwen3 technical report, 2025

  21. [29]

    Qwen2.5 technical report, 2025

    Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, ...

  22. [30]

    Gemma 3 technical report

    Gemma Team . Gemma 3 technical report. CoRR, abs/2503.19786, 2025

  23. [31]

    The llama 3 herd of models, 2024

    Llama3 Team . The llama 3 herd of models, 2024

  24. [32]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...

  25. [33]

    Measuring massive multitask language understanding, 2021 a

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding, 2021 a

  26. [34]

    Measuring mathematical problem solving with the math dataset, 2021 b

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset, 2021 b

  27. [35]

    Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023

    Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023

  28. [36]

    Language models are multilingual chain-of-thought reasoners, 2022

    Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. Language models are multilingual chain-of-thought reasoners, 2022

  29. [37]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer, 2023

  30. [38]

    Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021

    Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021

  31. [39]

    Pointer sentinel mixture models, 2016

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models, 2016

  32. [40]

    Openmathinstruct-2: Accelerating ai for math with massive open-source instruction data, 2024

    Shubham Toshniwal, Wei Du, Ivan Moshkov, Branislav Kisacanin, Alexan Ayrapetyan, and Igor Gitman. Openmathinstruct-2: Accelerating ai for math with massive open-source instruction data, 2024

  33. [41]

    Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms, 2024

    Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms, 2024

  34. [42]

    Training verifiers to solve math word problems, 2021

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021

  35. [43]

    O lympiad B ench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems

    Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. O lympiad B ench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientifi...

  36. [44]

    AIME 2024 , 2024

    HuggingFaceH4 . AIME 2024 , 2024. Dataset, accessed 2025-11-21

  37. [45]

    Gshard: Scaling giant models with conditional computation and automatic sharding, 2020

    Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. Gshard: Scaling giant models with conditional computation and automatic sharding, 2020

  38. [46]

    Emergent modularity in pre-trained transformers, 2023

    Zhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Chaojun Xiao, Xiaozhi Wang, Xu Han, Zhiyuan Liu, Ruobing Xie, Maosong Sun, and Jie Zhou. Emergent modularity in pre-trained transformers, 2023

  39. [47]

    Smith, and Luke Zettlemoyer

    Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, and Luke Zettlemoyer. Demix layers: Disentangling domains for modular language modeling, 2021

  40. [48]

    Are neural nets modular? inspecting functional modularity through differentiable weight masks, 2021

    Róbert Csordás, Sjoerd van Steenkiste, and Jürgen Schmidhuber. Are neural nets modular? inspecting functional modularity through differentiable weight masks, 2021

  41. [49]

    Recurrent independent mechanisms, 2020

    Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf. Recurrent independent mechanisms, 2020

  42. [50]

    Knowledge neurons in pretrained transformers, 2022

    Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. Knowledge neurons in pretrained transformers, 2022

  43. [51]

    Transformer feed-forward layers are key-value memories, 2021

    Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. Transformer feed-forward layers are key-value memories, 2021

  44. [52]

    Locating and editing factual associations in gpt, 2023

    Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in gpt, 2023

  45. [53]

    The lottery ticket hypothesis for pre-trained bert networks, 2020

    Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin. The lottery ticket hypothesis for pre-trained bert networks, 2020

Pith tools

Reviewed June 30, 2026 · model on record in the stance chip above.