{"id":"d287d67f-7145-4909-b959-0d49ff0539e3","arxiv_id":"2605.24846","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Keystone neurons are a stable, sparse, pretraining-established subset in Transformers; their removal collapses behavior and updating only them yields comparable task gains to full fine-tuning.","lead":"The paper identifies a sparse subset of 'keystone neurons' in LLMs that stay highly activated across many tasks; removing them collapses model behavior. It shows that fine-tuning only these neurons can match or beat full-parameter updates while using far fewer parameters and preserving other skills.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Performance collapse after removing activation-selected neurons may reflect selection artifact rather than unique criticality of the subset","rationale":"The reader's weakest_assumption directly isolates the same selection-bias risk that is load-bearing for the causal claim. Because the provided abstract supplies no ablation controls or layer-wise statistics that would refute the artifact interpretation, the concern remains unresolved even after reading the full manuscript description. No independent verification (reproducible code, formal proof) is mentioned that would bypass the need for this control experiment.","tokens_in":1762,"tokens_out":344,"duration_ms":12916,"concrete_test":"Ablate an equal number of neurons chosen uniformly at random within the same layers as the reported keystone set (or chosen by per-task activation rank) and recompute the capability-collapse metrics; if the magnitude of degradation is statistically indistinguishable from the keystone ablation, the causal attribution to the cross-task selection criterion does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the identified keystone neurons are causally responsible for capability collapse upon removal, and that their parameters are specially calibrated during pretraining. The isolation procedure selects on cross-task activation strength; this creates a potential circularity where any sufficiently high-activation subset (or any subset correlated with high-magnitude weights) would produce similar collapse when ablated. Without an explicit control that matches the selection criteria (e.g., same cardinality, same layer distribution) but uses an orthogonal criterion such as random sampling, single-task activation, or gradient-based importance, the observed drop cannot be attributed specifically to the cross-task consistency property rather than to generic properties of high-activation neurons.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript claims that across open-weight Transformers, a sparse subset of neurons exhibits consistently high activation across tasks spanning multiple capability dimensions. By selecting on cross-task activation strength, an extremely sparse 'keystone' subset is isolated whose removal produces behavioral collapse; these neurons are argued to be stable, intrinsic, and largely fixed during pretraining with tightly calibrated parameters. The work further proposes a supervised fine-tuning procedure that updates only the keystone neurons and reports task gains comparable or superior to full-parameter fine-tuning while better preserving performance on unrelated capabilities despite modifying far fewer parameters.","tokens_in":1912,"tokens_out":593,"duration_ms":26419,"significance":"If the empirical claims are substantiated with appropriate controls and quantitative results, the identification of a stable, pretraining-established sparse subset whose targeted update yields efficient adaptation would be a notable contribution to LLM interpretability and parameter-efficient fine-tuning. The potential reduction in modified parameters while maintaining or improving multi-task performance could influence both mechanistic understanding and practical deployment.","major_comments":[{"comment":"Abstract and method description: the central claim that ablation of the cross-task high-activation subset produces collapse specifically because of the cross-task consistency property (rather than generic high-activation or magnitude properties) requires explicit controls. No description is given of matched-cardinality random subsets, single-task activation subsets, or gradient-based importance baselines that would isolate the selection criterion; without these, the observed drop cannot be attributed to the stated mechanism.","section":"Abstract / probing procedure"},{"comment":"Abstract and experiments: the fine-tuning claim that updating only keystone neurons yields 'comparable or even better' task gains while better preserving other capabilities is presented without any reported metrics, number of updated parameters, task suite, baseline comparisons, or ablation on the selection threshold. These quantitative details are load-bearing for the practical contribution and are absent from the provided text.","section":"Abstract / fine-tuning section"},{"comment":"Abstract: the assertion that keystone neurons are 'largely established during pretraining' and that their 'precise values are critical' is stated without supporting evidence such as activation statistics across training checkpoints, parameter-sensitivity analysis, or comparison to randomly initialized models. This is central to the intrinsic-property claim yet unsupported in the given material.","section":"Abstract / analysis of pretraining stability"}],"minor_comments":[{"comment":"Notation for activation strength and the precise definition of the 'probing along the cross-task activation strength' procedure should be formalized with an equation or algorithm box to allow replication.","section":"Method"},{"comment":"The term 'keystone neurons' is introduced without reference to prior related concepts in the interpretability literature (e.g., 'critical neurons' or 'superposition' studies); a brief related-work paragraph would clarify novelty.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The provided materials contain only the abstract and no quantitative tables, figures, or experimental details, which prevents verification of any of the stated empirical claims. This raises a scope concern for a methods-oriented venue until the full experimental section is supplied."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their valuable feedback on our manuscript. We address each of the major comments point by point below, and we will make the necessary revisions to strengthen the empirical support for our claims.","responses":[{"response":"We agree that the manuscript requires these controls to properly attribute the effect to the cross-task consistency. The current version does not provide descriptions of matched-cardinality random subsets, single-task activation subsets, or gradient-based baselines. We will add these controls to the probing procedure and results in the revised manuscript.","revision_made":"yes","referee_comment":"[Abstract / probing procedure] Abstract and method description: the central claim that ablation of the cross-task high-activation subset produces collapse specifically because of the cross-task consistency property (rather than generic high-activation or magnitude properties) requires explicit controls. No description is given of matched-cardinality random subsets, single-task activation subsets, or gradient-based importance baselines that would isolate the selection criterion; without these, the observed drop cannot be attributed to the stated mechanism."},{"response":"The referee is correct that the provided manuscript text does not include the specific quantitative metrics, parameter counts, task suite details, baseline comparisons, or threshold ablations for the fine-tuning experiments. We will revise the experiments section to report these details comprehensively, including a table with the metrics and ablations.","revision_made":"yes","referee_comment":"[Abstract / fine-tuning section] Abstract and experiments: the fine-tuning claim that updating only keystone neurons yields 'comparable or even better' task gains while better preserving other capabilities is presented without any reported metrics, number of updated parameters, task suite, baseline comparisons, or ablation on the selection threshold. These quantitative details are load-bearing for the practical contribution and are absent from the provided text."},{"response":"We acknowledge that the current manuscript does not include the supporting evidence such as activation statistics across checkpoints, parameter-sensitivity analysis, or comparisons to random initialization. We will perform and incorporate these analyses into the revised version to support the claims about pretraining stability and parameter criticality.","revision_made":"yes","referee_comment":"[Abstract / analysis of pretraining stability] Abstract: the assertion that keystone neurons are 'largely established during pretraining' and that their 'precise values are critical' is stated without supporting evidence such as activation statistics across training checkpoints, parameter-sensitivity analysis, or comparison to randomly initialized models. This is central to the intrinsic-property claim yet unsupported in the given material."}],"tokens_in":1441,"tokens_out":549,"duration_ms":55102,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that the authors locate a very small set of neurons that stay highly active across many tasks, show that ablating them tanks performance, and then demonstrate that fine-tuning only those neurons can deliver task gains comparable to updating everything while hurting other capabilities less.\n\nWhat is new is the specific pipeline: probe activation strength across a range of capability dimensions to pick the subset, confirm it is stable from pretraining, and restrict supervised fine-tuning to just those parameters. The abstract indicates this holds on multiple open-weight Transformers and that the parameters are tightly calibrated early.\n\nThe practical angle is useful if the numbers check out. Restricting updates to a tiny fraction of weights while preserving breadth is the kind of result people working on efficient adaptation would notice.\n\nThe soft spot is the missing controls on the selection step itself. Picking on cross-task activation and then observing collapse does not yet rule out that any high-magnitude or high-activation subset of the same size would produce the same drop. The stress-test concern lands: without an orthogonal match (random, single-task, or gradient-based) at the same cardinality and layer distribution, the claim that these neurons are specially critical because of their cross-task consistency remains unproven. The abstract also gives no concrete sparsity figures, task lists, or metric tables, so the strength of the empirical support cannot be judged from what is shown.\n\nThis is for researchers already thinking about neuron-level interventions and parameter-efficient tuning. A reader looking for a new lens on what makes models capable could pull ideas from it, but the current evidence is not yet tight enough for strong claims.\n\nI would send it to peer review. The direction is worth a proper referee look even if the controls and numbers need tightening.","headline":"The paper isolates a sparse cross-task activation subset whose removal collapses behavior and whose selective fine-tuning matches full updates, but the causal link to unique criticality still needs matched controls.","tokens_in":2398,"tokens_out":435,"would_cite":false,"duration_ms":20014,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A sparse set of keystone neurons, isolated by consistent high activation across tasks, controls core capabilities in open-weight Transformers.","keywords":["keystone neurons","LLM interpretability","neuron activation","sparse subsets","fine-tuning","transformer models","pretraining","capability preservation"],"falsifier":"An experiment in which randomly chosen neurons matched for activation strength produce the same collapse upon removal, or in which keystone neurons selected from one prompt set fail to affect held-out tasks.","tokens_in":2674,"feed_emoji":"🧠","tokens_out":660,"duration_ms":35953,"temperature":0.7,"pith_summary":"The paper examines neuron activations inside large language models during inference on varied tasks. It identifies an extremely small subset that stays highly active no matter the task and shows that deleting this subset destroys model performance. These keystone neurons form early in pretraining and stay fixed, with their exact parameter values proving essential. The authors then demonstrate that fine-tuning only this subset produces gains equal to or better than updating every parameter while harming unrelated abilities less. A sympathetic reader would see this as evidence that models rely on a tiny, stable core rather than distributed computation across all neurons.","feed_headline":"Sparse keystone neurons collapse LLM performance when removed","feed_subtitle":"These neurons, fixed early in pretraining, let targeted fine-tuning match full updates while sparing other abilities","key_machinery":"Keystone neurons: the sparse subset isolated by high cross-task activation strength whose removal collapses model behavior.","core_discovery":"Across a wide range of open-weight Transformers, a subset of neurons remains consistently highly activated during inference across tasks of multiple capability dimensions. By probing along the cross-task activation strength, an extremely sparse subset is isolated, whose removal causes a collapse in model behavior, which we term keystone neurons. Our analysis reveals that keystone neurons are a stable and intrinsic neuron subset of the model that is largely established during pretraining. The parameters associated with these neurons are tightly calibrated during the training process, and their precise values are critical for the capabilities of the model.","pith_inferences":["The finding could enable prompt-only auditing of model internals without full retraining.","It raises the possibility that similar sparse critical subsets exist in non-Transformer architectures.","Targeted editing of keystone neurons might support more precise capability addition or removal.","The approach could extend to measuring how pretraining data distributions shape these stable subsets."],"forward_implications":["Updating only keystone neurons during supervised fine-tuning yields task gains comparable to or better than full-parameter fine-tuning.","Targeted updates on keystone neurons better preserve performance in other capability dimensions.","Keystone neurons form a stable intrinsic subset largely established during pretraining.","Precise parameter values tied to these neurons are critical for overall model capabilities."],"fun_headline_variants":["Removing keystone neurons collapses LLM performance","LLM keystone neurons fixed during pretraining","Keystone fine-tuning matches full parameter updates","Activation probing reveals keystone neurons"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The high cross-task activation and the performance collapse after removal are caused by these neurons being uniquely critical rather than by the selection method itself or by other components that were not isolated.","fun_headline_variants_meta":{"raw":{"variants":["Removing keystone neurons collapses LLM performance","LLM keystone neurons fixed during pretraining","Keystone fine-tuning matches full parameter updates","Activation probing reveals keystone neurons"]},"model":"grok-4.3","cost_usd":0.005867,"raw_usage":{"total_tokens":2701,"prompt_tokens":655,"num_sources_used":0,"completion_tokens":51,"cost_in_usd_ticks":58665500,"prompt_tokens_details":{"text_tokens":655,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1995,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":655,"tokens_out":51,"duration_ms":18215,"temperature":1.0,"reasoning_tokens":1995,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T12:43:45.771537+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An experiment in which randomly chosen neurons matched for activation strength produce the same collapse upon removal, or in which keystone neurons selected from one prompt set fail to affect held-out tasks.","supporting_citations":[],"review_version":1}