{"id":"00412a74-2a53-40ae-9ea3-6650f8263029","arxiv_id":"2501.08618","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Pretrained LLMs use largely disjoint neuron sets for hierarchical versus linear grammaticality judgments, and hierarchical-selective neurons transfer to nonsense-word grammars.","lead":"This paper asks whether large language models develop separate internal machinery for hierarchical versus linear grammar rules, echoing brain studies in humans. It finds that the neurons most involved in judging hierarchical sentences differ from those for linear rules, and that the separation persists even for invented nonsense words.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The hierarchical stimuli differ from linear stimuli in naturalness, not only in structural type: hierarchical positives are natural sentences and negatives are scrambled, while linear positives and negatives are both artificial; the Jabberwocky control does not remove this confound, so disjoint…","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the grammar design confounds structural type with naturalness. I agree that this is the most serious threat to the central claim because it undermines the interpretation of every experiment, including the Jabberwocky control. The paper's footnote 9 concedes the behavioral results could be explained by low-probability task/inputs, but the component localizations are presented as structural evidence; the confound applies to them equally. Other issues, such as the non-significant linear-side ablation comparisons in Table 9, are important but secondary: even if the linear-side causal claim were fixed, the naturalness confound would still leave the hierarchy-versus-linearity interpretation ambiguous. The proposed likelihood-matching test is a concrete, decisive check because it directly measures the naturalness/in-distributionness dimension that the stimulus design fails to hold fixed. If the effects survive likelihood matching, the authors' interpretation is strongly supported; if not, the evidence reduces to models distinguishing natural from unnatural word order. This does not change the reader's conditional verdict, which already requires additional controls and honest re-reporting.","tokens_in":23104,"tokens_out":6298,"duration_ms":70242,"concrete_test":"Compute the per-stimulus log-likelihood (or pseudo-log-likelihood) under each base model for every positive and negative example in all grammars, including the Jabberwocky sets. Then re-run Experiment 1's accuracy comparison and Experiment 2's component-overlap analysis on likelihood-matched subsets: match each hierarchical positive-negative pair to a linear pair with a comparable log-likelihood difference, or include the log-likelihood difference as a covariate in the statistical tests. If the hierarchical-versus-linear accuracy gap and the H/L overlap differences disappear or become non-significant after matching, the results are explained by in-distributionness/naturalness rather than structural type; if they persist on matched items, the hierarchy/linearity interpretation is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that hierarchical and linear grammars recruit disjoint mechanisms rests on stimuli that conflate structural type with naturalness. In Section 2.2, every hierarchical negative is formed by swapping the final two words of a positive example, so hierarchical positives are natural, in-distribution sentences and hierarchical negatives are scrambled. In contrast, linear positives are generated by positional rules such as inserting a word at position 5, so both linear positives and negatives are artificial and out-of-distribution. The model could therefore solve the grammaticality judgment by detecting whether the word order is natural (hierarchical condition) versus following an arbitrary positional pattern (linear condition), without any representation of hierarchy per se. The authors acknowledge this behavioral confound in footnote 9, but the component localization and ablation results inherit the same confound: the top-1% neurons identified for hierarchical grammars may be 'natural sentence' detectors, and those for linear grammars may be 'artificial pattern' detectors, yielding low overlap and selective ablation effects that have nothing to do with hierarchy. Experiment 4 attempts to control for lexical in-distributionness using Jabberwocky words, but it preserves the same construction: hierarchical Jabberwocky positives follow natural subject-verb-object order and negatives are formed by swapping the final two words, while linear Jabberwocky positives still contain position-based insertions. Thus the nonce control removes lexical familiarity but not the natural-versus-scrambled word-order asymmetry. If the observed behavioral gap, component disjointness, and ablation selectivity are driven by this naturalness difference rather than by hierarchical versus linear structure, the paper's central claim is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether pretrained LLMs develop separable mechanistic resources for processing hierarchical versus linear/positional grammatical structure, inspired by Musso et al. (2003). It constructs three hierarchical and three linear grammars in English, Italian, Japanese, and nonce versions, evaluates six LLMs on in-context grammaticality judgments, localizes the top 1% of attention/MLP units by attribution patching, measures pairwise overlap between grammars, and tests causal specificity through ablations. The authors report higher accuracy on hierarchical than linear grammars for English and Italian, greater component overlap within hierarchical grammars than across hierarchical-linear pairs, and selective ablation effects for hierarchical components, including on Jabberwocky stimuli; they conclude that LLMs acquire localizable, largely disjoint processing mechanisms for hierarchical versus linear structure from distributional exposure alone.","tokens_in":23393,"tokens_out":6893,"duration_ms":70534,"significance":"If the conclusions are valid, the paper would be an important contribution to mechanistic interpretability and to debates about inductive biases in language models: it would show functional specialization for syntax-like structure arising in general-purpose sequence models without explicit human-like biases. The study has real strengths: it spans six open-weight models, three natural languages plus Jabberwocky controls; attribution patching makes component localization computationally tractable; ablations go beyond correlational overlap; the authors state that data and code are released; and the Limitations section candidly discusses polysemanticity, task specificity, and an alternative behavioral explanation. However, the stimulus design conflates hierarchy with natural word order, and the ablation statistics in Table 9 do not support linear-selective components. These issues are central rather than cosmetic, so the current manuscript cannot be accepted as-is.","major_comments":[{"comment":"The stimulus construction conflates grammatical type with naturalness. For every hierarchical grammar, positives are ordinary natural-language sentences and negatives are formed by swapping the final two words; for every linear grammar, both positives and negatives are arbitrary positional manipulations (e.g., inserting a word at position 4 or reversing the sentence). Therefore higher accuracy on hierarchical grammars, and the top-1% neurons identified for them, could reflect detection of 'natural word order' versus 'artificial word order' rather than hierarchy per se. Experiment 4 replaces the lexicon with Jabberwocky words but preserves exactly the same asymmetry: hierarchical Jabberwocky positives follow natural SVO/SOV order, hierarchical negatives are final-two-word swaps, and linear items are positional patterns. Hence the Jabberwocky control removes only lexical in-distributionness, not the naturalness confound. Footnote 9 acknowledges this alternative for behavioral results, but the same alternative threatens Experiments 2 and 3. A concrete control would be to generate hierarchical negatives by moving a word to an arbitrary non-final position, or to match hierarchical and linear items on surface plausibility, and then re-run the overlap and ablation analyses.","section":"§2.2 and §3.4"},{"comment":"The causal claim for linear-selective components is not supported by the reported statistics. Table 9 shows that on linear grammars, ablating H-components versus L-components is non-significant in every language: EN(L) p=0.73, IT(L) p=0.33, JP(L) p=0.06. Appendix B.3 also states that 'relative accuracy decreases are not significantly different between ablations of hierarchical/linear components' for linear grammars. The main text says 'Ablating components from L decreases the model's accuracy on linear structures more than ablating H,' which contradicts these data. Since the disjointness claim requires selectivity in both directions, the paper currently supports only hierarchical-selective components, and that support is largely relative to random ablation rather than to L-components.","section":"§3.3 / Table 9"},{"comment":"The behavioral claim is language-dependent, but the paper's framing overgeneralizes. For Japanese, the Mann-Whitney U test in Table 4 is U=203, p=0.2, so there is no significant accuracy advantage for hierarchical over linear grammars in Japanese. The text in §3.1 carefully limits the p<0.001 statement to English and Italian, yet RQ1 and the Conclusion say that models 'show distinct behaviors' on hierarchical versus linear inputs without the Japanese qualifier. The authors should either state RQ1 as restricted to English and Italian or explain why the Japanese null does not undermine the cross-linguistic claim.","section":"§3.1 / Table 4"},{"comment":"The interpretation of the overlap results needs a chance baseline. The paper reports that all pairwise overlaps are significantly different from zero and that H-H overlaps exceed H-L overlaps. However, because top-1% sets are selected using the same task format and answer tokens, a non-zero baseline overlap is expected even for fully shared mechanisms; the H-H vs H-L contrast is the relevant comparison, and the significance there is encouraging. Still, the statement that overlaps are 'significantly different from 0' is not evidence of specialization. A permutation baseline or a random-subsample overlap calculation should be reported so that the reader can see how much of the absolute overlap is expected by chance.","section":"§3.2 and §3.4"}],"minor_comments":[{"comment":"The model list contains six models, but Appendix Tables 7, 8, and 12 refer to '7 models' or report N=108 without explaining the count; please reconcile the model count and the sample-size notation.","section":"§2.1"},{"comment":"The negative examples in Table 1 are missing spaces (e.g., 'a woman readschaptera'), which is presumably a typesetting artifact but should be corrected for readability.","section":"Table 1"},{"comment":"The github link is given only as 'github'; a proper URL is needed to substantiate the reproducibility claim.","section":"§2.2"},{"comment":"The Limitations section discusses polysemanticity and task generalization but does not list the naturalness confound of the stimulus construction as a limitation; given its central role, it should be acknowledged there.","section":"Limitations"},{"comment":"Labels such as 'H(ZZ x EN)' are not defined in the main text; please add a sentence explaining this cross-grammar component-overlap notation.","section":"§3.4 / Table 12"},{"comment":"The prose around Figure 4d should be more careful: Table 11 shows a significant effect of English hierarchical ablations on Jabberwocky hierarchical accuracy (p=0.004) but not on Jabberwocky linear accuracy (p=0.41), so the claim should be restricted to the hierarchical condition.","section":"§3.4 / Table 11"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid scaffolding and an honest limitations section, but the two main technical objections—the naturalness confound and the non-significant linear ablation—cut to the central 'disjoint mechanisms' claim. I think the authors can address this within a revision by adding a naturalness-controlled stimulus set, re-running the overlap and ablation analyses, and substantially softening the two-way disjointness interpretation. If the revised analyses still show H-H > H-L overlap and H-selective ablation under naturalness controls, the paper would be a strong contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuine contribution to mechanistic interpretability, but the central claim is overweighted. The overlap results are solid — hierarchy-selective components cluster together and are distinct from linear-selective components across languages. The causal story is lopsided: ablating hierarchy components significantly hurts hierarchy judgments, but ablating linear components does not selectively hurt linear judgments (Table 9: p = 0.73, 0.33, 0.06 for EN/IT/JP). So the paper demonstrates a hierarchy-selective mechanism, not a disjoint pair of mechanisms.\n\nWhat's new: this is the first causal localization of a hierarchical-versus-linear distinction in pretrained LLMs, including transfer to nonce grammars. The design is genuinely careful: six open-weight models, three languages, attribution patching, and ablation baselines. The authors are also honest — footnote 9 flags the teleological concern, and the appendix reports the non-significant linear ablations.\n\nThe main soft spot is stimulus construction. Hierarchical positives are natural sentences; hierarchical negatives are formed by swapping the final two words, producing scrambled word order. Linear positives are themselves artificial (e.g., insert 'doesn't' at position 5). The model could be detecting natural word order versus arbitrary positional patterns, with no representation of hierarchy per se. The Jabberwocky control does not remove this: hierarchical Jabberwocky still has SVO order versus a swapped final pair, while linear Jabberwocky still uses position-based insertions. The nonce experiment strips lexical familiarity but preserves the naturalness asymmetry. So the confound affects the component localizations as much as the behavioral accuracy gap.\n\nTwo smaller issues: the statistical appendix contains suspicious copy-paste — Table 8 gives identical test statistics and p-values for English, Italian, and Japanese. And §2.2 says data/code are on GitHub, but no link appears in the text.\n\nWhere this leaves the paper: the hierarchy-selective components are probably real, but 'disjoint processing mechanisms for hierarchical and linear grammars' is not established. The supported claim is weaker: LLMs have components that are causally important for judging naturalistic sentences, and these differ from components used for artificial positional patterns. That is still interesting, but it is not the brain-inspired claim.\n\nMy recommendation: send it to peer review. The question is important, the methods are sophisticated, and the flaws are addressable. A good referee will ask for matched naturalness controls, a proper re-analysis of the linear ablations, and code/data. I would not cite the central claim as established, but I would cite the overlap result. Worth a reading group as a case study in stimulus confounds.","headline":"A serious localization study whose central disjointness claim is undermined by a stimulus confound and by non-significant linear-side ablations.","tokens_in":23959,"tokens_out":3665,"would_cite":false,"duration_ms":34383,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T50"],"pacs":[],"model":"deepseek-v4-flash","headline":"Pretrained LLMs process hierarchical and linear grammars with largely separate internal components.","keywords":["hierarchical grammar","linear grammar","functional specialization","mechanistic interpretability","attribution patching","grammaticality judgment","large language models","Jabberwocky sentences"],"falsifier":"Construct a version of the experiment where linear positives are as natural-sounding as hierarchical positives, using attested word-order variants, while hierarchical negatives are made exactly as unnatural as the linear negatives; if the accuracy gap and the hierarchical-versus-linear component disjointness then vanish, the central claim would be falsified.","tokens_in":22909,"feed_emoji":"🧠","tokens_out":6678,"duration_ms":66823,"temperature":0.7,"pith_summary":"The paper asks whether a general-purpose learner exposed only to text can acquire the kind of functional segregation that human brains show for hierarchical language structure. Using synthetic grammars in English, Italian, Japanese, and nonce words, it first shows that pretrained LLMs are more accurate at grammaticality judgments for hierarchical than for linear/positional rules. It then locates the top 1% of MLP and attention neurons most influential for each judgment via attribution patching, and reports that hierarchy-selective neurons overlap strongly across hierarchical grammars but only weakly with linear-selective neurons; ablating them selectively degrades hierarchical judgments. Hierarchy-selective components also drive judgments on meaningless Jabberwocky sentences. The paper concludes that functional specialization toward hierarchical syntax can emerge from distributional exposure alone, and that the responsible components are localizable and partially abstract.","feed_headline":"LLMs use separate circuits for hierarchical vs linear grammar","feed_subtitle":"Neuron-level ablations show hierarchy-sensitive components transfer across languages and nonsense words.","key_machinery":"The central object is a controlled grammar battery: eighteen synthetic grammars, with three hierarchical structures (declarative, subordinate, passive) and three linear structures (negation, inversion, and a language-specific third rule such as wh-word insertion, last-noun agreement, or past-tense placement), each generated in English, Italian, Japanese, and again with nonce words. On each grammar the model performs an in-context grammaticality judgment, and component importance is scored by attribution patching, a first-order Taylor approximation of the indirect effect of each MLP and attention neuron on the logit difference between the correct and incorrect answer tokens. The top 1% of neurons by estimated effect define the hierarchy-sensitive set H and the linearity-sensitive set L; pairwise overlap between these sets and mean-activation ablation experiments test whether the sets are causally distinct.","core_discovery":"On the paper's own terms, the discovery is that pretrained large language models contain largely disjoint, causally verified component sets for processing hierarchical versus linear grammars. Across six open-weight models, the top 1% of neurons by estimated indirect effect on hierarchical grammaticality judgments show high pairwise overlap within the hierarchical family and across English, Italian, and Japanese, but significantly lower overlap with the top 1% for linear grammars. Ablating the hierarchy-sensitive set to its mean activation reduces accuracy on hierarchical inputs more than ablating the linear-sensitive set or a random set, and the reverse holds for linear inputs. The same hierarchy-sensitive components also influence grammaticality judgments on nonce-word sentences, which the paper takes as evidence that the specialization tracks abstract structure rather than lexical meaning or training-distribution familiarity.","pith_inferences":["A natural next experiment would balance word-order naturalness directly: construct linear-positive sentences using word orders that actually occur in some dialect or register, and hierarchical negatives using swaps that preserve attested word order; if the accuracy gap and component disjointness then vanish, the results would be explained by natural-versus-artificial order rather than hierarchy-ve","The same attribution-patching protocol could test other structural contrasts, such as center-embedding, cross-serial dependencies, or artificial grammars matched for surprisal, to see whether hierarchy is the operative variable or a proxy for some other distributional property.","Because the analysis operates on individual neurons and attention outputs, future work with more fine-grained units might reveal that the 'disjoint' sets are not separate mechanisms but different overlapping multifunctional circuits; the localization claim may be coarser than the true causal structure.","If confirmed, the finding suggests LLM pretraining could serve as a controlled testbed for how functional specialization in language arises from exposure, complementing human neuroimaging studies where such controlled exposure is impossible."],"forward_implications":["If correct, the paper implies that pretrained LLMs develop localizable subnetworks specialized for hierarchical syntax, so syntax processing is not smeared uniformly across the network.","Hierarchy-selective components are shared across languages and transfer to nonce inputs, implying the specialization is a structural abstraction rather than a lexicon effect.","The selective ablation results give causal, not merely correlational, evidence that these component sets are functionally distinct.","The selectivity is not uniform: smaller models and Japanese inputs show weaker, less selective effects, indicating that specialization strengthens with model scale and with language prevalence in the training distribution."],"supporting_citations":[{"why":"Supplies the human experiment the paper replicates: identical vocabularies with only hierarchical grammars engaging language areas.","marker":"Musso et al. (2003)"},{"why":"Shows autoregressive transformers learn hierarchical grammars more easily, which motivates the behavioral expectation tested in pretrained models.","marker":"Kallini et al. (2024)"},{"why":"Establishes that human language circuits respond less to Jabberwocky sentences, the benchmark for the paper's nonce-word experiments.","marker":"Fedorenko et al. (2016)"},{"why":"Provides the attribution-patching estimator used to score every MLP and attention neuron in the model.","marker":"Kramár et al. (2024)"},{"why":"Validates attribution patching as a scalable alternative to activation patching for localizing model behavior.","marker":"Syed et al. (2024)"},{"why":"Formulates the teleological alternative that low-probability linear inputs are harder regardless of mechanism, which the paper cites as a confound.","marker":"McCoy et al. (2024)"},{"why":"Prior evidence that LLMs contain causally task-relevant language-selective units, which the paper's causal localizations extend.","marker":"AlKhamissi et al. (2024)"},{"why":"Prior evidence of brain-like functional organization in LLMs, supporting the idea of syntax-selective subnetworks.","marker":"Sun et al. (2024)"}],"fun_headline_variants":["LLMs use separate circuits for hierarchical and linear grammar","Causal ablation shows distinct LLM circuits for grammar types","Hierarchy and linear grammar spark separate transferable LLM circuits","LLMs split grammar processing into disjoint hierarchical and linear pathways"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the only systematic difference between hierarchical and linear stimuli is structural type; but because hierarchical positives are natural word-order sentences and hierarchical negatives are made by swapping the final two words, the models might be distinguishing natural from scrambled word order rather than hierarchy from linearity, a confound the paper acknowledges only in a footnote.","fun_headline_variants_meta":{"raw":{"variants":["LLMs use separate circuits for hierarchical and linear grammar","Causal ablation shows distinct LLM circuits for grammar types","Hierarchy and linear grammar spark separate transferable LLM circuits","LLMs split grammar processing into disjoint hierarchical and linear pathways"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1168,"prompt_tokens":860,"completion_tokens":308,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":241}},"tokens_in":476,"tokens_out":308,"duration_ms":3543,"temperature":1.0,"reasoning_tokens":241,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:21:41.553906+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a version of the experiment where linear positives are as natural-sounding as hierarchical positives, using attested word-order variants, while hierarchical negatives are made exactly as unnatural as the linear negatives; if the accuracy gap and the hierarchical-versus-linear component disjointness then vanish, the central claim would be falsified.","supporting_citations":[{"cited_title":"u rgen Reichenbach, Christian B \\","cited_arxiv_id":null,"evidence_quote":"Supplies the human experiment the paper replicates: identical vocabularies with only hierarchical grammars engaging language areas."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Validates attribution patching as a scalable alternative to activation patching for localizing model behavior."},{"cited_title":"Brain-like Functional Organization within Large Language Models","cited_arxiv_id":"2410.19542","evidence_quote":"Prior evidence of brain-like functional organization in LLMs, supporting the idea of syntax-selective subnetworks."}],"review_version":1}