{"id":"1e0c2e27-0b1d-4210-921c-cd93b5224984","arxiv_id":"2601.21587","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Bilingual GPT-2 models show that later L2 introduction strengthens the correlation between syntactic distance and crosslinguistic interference, though the claimed causal attention ablations are missing from the paper.","lead":"This paper trains bilingual language models that learn English after one of five other languages and measures how the first language interferes with the second. It finds that later exposure to English makes interference track linguistic distance, but the strongest mechanistic claims in the abstract are not backed by the experiments shown.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's central mechanistic claim—'deep-layer attention mechanisms exclusively drive' CLI—has no corresponding experiment in the paper; §5 and appendices contain only behavioral ablations, so the headline causal claim is unsupported.","rationale":"The central claim has two parts: a behavioral association between CLI and syntactic distance modulated by age of exposure, and a mechanistic claim that deep-layer attention exclusively drives transfer. The behavioral part has some support: a controlled age-of-exposure manipulation, per-pair tokenizers, an early-imbalance control, and distance-consistent correlations over five languages. But the mechanistic exclusivity claim requires a causal manipulation of attention, and no such experiment is reported. The abstract explicitly promises 'targeted causal ablations' and 'exclusively drive'; the body's Section 5 and Appendix B never deliver this. This is not a stylistic mismatch: it is the difference between a correlational finding and the paper's advertised contribution. The reader's weakest_assumption is the WALS-based syntactic distance and NLLB prime quality; those are legitimate concerns for the behavioral correlation. My concern is different and more directly decisive: even if the distance measure and translations were perfect, the paper would still not support the exclusivity claim because the relevant ablation is absent. I therefore agree with the REJECT verdict, but for a partly different reason. If the authors supply the missing attention-ablation experiment and it isolates deep-layer attention, the mechanistic portion could be reassessed; as submitted, the central claim is unsupported.","tokens_in":13484,"tokens_out":5158,"duration_ms":50069,"concrete_test":"Locate or reproduce the promised causal ablation: on the Step=16K/32K models with German and Turkish L1s, ablate all attention heads in layers 9–12 during primed BLiMP evaluation and compare CLI against (a) unablated priming, (b) ablation of the corresponding MLP blocks, and (c) ablation of lower-layer attention. If distance-consistent priming survives deep-layer attention ablation, the exclusivity claim is false; if it disappears specifically with deep attention but not with MLP or shallow-layer ablation, the claim gains support. The current manuscript includes no such condition, so this experiment would settle whether the advertised mechanism exists.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central mechanistic claim, stated in the abstract, is that 'targeted causal ablations ... confirm that deep-layer attention mechanisms exclusively drive this crosslinguistic transfer.' The manuscript body does not contain that experiment. Section 5 reports three ablations—early imbalanced training (§5.1), L1/L2 reversal (§5.2), and model scale (§5.3)—none of which involve attention. Appendix B contains priming-by-phenomena, translation-sensitivity, and permutation figures, again with no attention manipulation. The only mechanistic evidence offered is correlational: LogitLens L1-token ratios (§4.3.1) and L2 neuron-overlap correlations (§4.3.2). Correlation cannot establish exclusivity, and neither analysis rules out alternative routes (MLPs, residual connections, lower layers, or tokenizer artifacts). Additionally, the abstract advertises '15 typologically diverse L1s,' while §3.2 trains 5. The behavioral core may survive, but the advertised mechanistic exclusivity claim is unsupported by the submitted evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains bilingual GPT-2 models on English and one of five L1s, varying the step at which L2 is introduced (SoE), and studies crosslinguistic influence (CLI) via BLiMP accuracy, FCE non-native preference, crosslinguistic priming, LogitLens, and L2 neuron overlap. The central claim is that later L2 introduction (higher L1 dominance, lower L2 proficiency) produces a syntactic-distance-dependent interference gradient, and that this gradient is mechanistically rooted in L1 co-activation and overlapping L2 neurons. The submitted body, however, does not contain the causal attention ablation advertised in the arXiv abstract, and the statistical support for the core correlation is incomplete.","tokens_in":13711,"tokens_out":7233,"duration_ms":82920,"significance":"If the behavioral findings hold, the paper would provide a controlled computational demonstration that age of exposure and typological distance jointly shape CLI, with potential implications for bilingualism theory and multilingual LM training. The use of external measures (BLiMP, FCE, WALS-based distances) and the attempt to connect behavior to internal representations are strengths. However, the advertised mechanistic exclusivity claim is unsupported, and the five-language, single-seed evidence base makes the central correlation fragile. The paper's contribution is therefore currently more modest than its abstract claims.","major_comments":[{"comment":"The abstract states that the study covers '15 typologically diverse L1s' and that 'targeted causal ablations ... confirm that deep-layer attention mechanisms exclusively drive' CLI. §3.2 trains only five L1s, and §5 reports no attention manipulation. The exclusivity claim is load-bearing and is not substantiated anywhere in the body or appendices. Either provide the attention ablation with appropriate controls, or remove the claim and reconcile the abstract with the actual scope.","section":"Abstract vs. §3.2, §5"},{"comment":"The central distance-gradient result rests on five languages and is plotted without error bars or confidence intervals. The text invokes permutation tests ('Appendix B.3') to support the result, but that appendix contains only Figure 12 and no test statistic, p-value, or permutation distribution. With five points, a single outlier can drive the correlation. Report the permutation results, add confidence intervals, and ideally increase the number of L1s or use multiple seeds.","section":"§4.1, Table 1, Appendix B.3"},{"comment":"The priming paradigm depends on NLLB translations preserving the target L1 structure, and the translation-quality ablation is presented only as two figures with no quantitative summary or description of the comparison. The text in §4.1 states that the ablation 'confirm[s]' results are not explained by lexical overlap or translation quality, but the submitted material does not support that assertion. Provide the actual comparison statistics and a narrative description of what was contrasted.","section":"§3.3.3, Appendix B.2"},{"comment":"The neuron-overlap result is reported as a 'consistent negative correlation' without a correlation coefficient, confidence interval, or sample size. The top-25% activation threshold and 'consistently ranking' criterion are not operationally defined. Since this is the main mechanistic evidence for the claim that 'L1 typological proximity physically dictates' L2 circuitry, it needs statistical rigor and a clearer description of the neuron identification procedure.","section":"§4.3.2, Figure 6"},{"comment":"The FCE preference measure normalizes by the model's own surprisal differences summed over all learner groups, so each model's baseline is confounded with its own L1. A control — such as a monolingual English model or a model without L1-specific training — is needed to demonstrate that the ΔS pattern reflects L1-specific preference rather than overall surprisal differences. Single-seed results (Appendix A, seed 123) also provide no uncertainty estimate.","section":"§3.3.2, Eq. (3)"}],"minor_comments":[{"comment":"Rename the 'Latin' column to 'Latin script' to avoid confusion with the Latin language family; the current columns 'Indo-European' and 'Latin' overlap in meaning.","section":"Table 1"},{"comment":"The notation Acc_{L1,L2} is ambiguous; clarify that it is the bilingual model's accuracy and explicitly state why a positive value is labeled 'interference' rather than 'negative transfer'.","section":"Eq. (2)"},{"comment":"There is a typo 'the the shared-syntax hypothesis', and the reference 'NICOLADIS (2006)' should be normalized to title case.","section":"§2, References"},{"comment":"The appendix headings are followed by figures but almost no prose. Add at least a paragraph per subsection describing the method, the metric, and the conclusion.","section":"Appendices B.1–B.3"},{"comment":"Add error bars or scatter distributions; the claimed 'wider confidence interval' for German is not visible from the current figure.","section":"Figure 6"},{"comment":"The full-text title differs from the arXiv title; align them in the final version.","section":"Title"}],"recommendation":"major_revision","confidential_remarks":"The abstract/body mismatch is the most serious issue: the advertised causal attention claim and the '15 L1s' figure are not supported by the submitted experiment. The paper's actual contribution is more modest and, with additional statistical detail and tightened claims, could be publishable. I recommend requiring the authors to either add the missing ablation or remove the mechanistic exclusivity claim, and to report the permutation/translation-ablation results numerically. The single-seed and five-language limitations should also be addressed explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: the paper's advertised contribution is two results at different strengths. The behavioral one—later L2 introduction strengthens the correlation between WALS-based syntactic distance and L2 interference on BLiMP—looks plausible and is worth testing further. The mechanistic one, as advertised in the abstract, is not in the paper. There is no attention ablation anywhere in Section 5 or the appendices. The abstract says 'targeted causal ablations ... confirm that deep-layer attention mechanisms exclusively drive this crosslinguistic transfer.' No such experiment exists. That's a load-bearing gap, not a cosmetic one.\n\nWhat's actually new: a clean-ish setup for varying age of exposure across five typologically distinct L1s, with a monolingual baseline that shares the bilingual tokenizer, and a metric for CLI that separates positive and negative transfer. The FCE non-native preference analysis, the LogitLens co-activation measure, and the neuron-overlap correlation are all reasonable exploratory instruments. The paper also engages the SLA-LM literature properly; the related-work discussion positions itself against Oba, Yadavalli, Constantinescu, and Arnett without hand-waving.\n\nSoft spots, in order of severity. First, the abstract overclaims on both the number of languages (15 vs. 5) and the attention ablation. That's a manuscript integrity issue that an editor should catch before review. Second, the mechanistic language—'physically dictates,' 'exclusively drive'—goes beyond correlation. The neuron-overlap correlation is interesting, but it doesn't rule out MLP or residual contributions or tokenizer artifacts. Third, the main Figure 2 has no error bars or confidence intervals; the permutation tests are referenced and Figure 12 shows something, but the main text doesn't describe the test procedure or how many permutations. That's fixable. Fourth, the WALS feature count as syntactic distance is a crude but defensible proxy; with only five languages, the correlation is fragile. The translations from NLLB are a genuine confound, though Figure 11 gives some reassurance.\n\nBottom line: the core behavioral finding could survive a serious referee, but the current manuscript doesn't earn the mechanistic headline. Worth engaging, but with the abstract heavily reframed and the missing analysis either supplied or retracted. I'd send it to review, but I'd expect major revision. For your own work, I wouldn't cite the causality claims, but the distance-by-soe interaction is potentially useful once it's backed by numbers.","headline":"Behavioral age-of-exposure effect is plausible, but the abstract promises a mechanistic attention ablation that the paper never runs—fix that mismatch before treating the causal claims seriously.","tokens_in":14237,"tokens_out":3474,"would_cite":false,"duration_ms":33630,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T50"],"pacs":[],"model":"deepseek-v4-flash","headline":"In bilingual language models, a dominant first language systematically warps the second language's grammar in proportion to syntactic distance.","keywords":["crosslinguistic influence","language models","bilingual training","structural priming","syntactic distance","language dominance","second language acquisition","neuron overlap"],"falsifier":"Train the same bilingual framework on a substantially larger set of L1s (e.g., 20 languages with densely sampled syntactic distances) and check whether the CLI-distance correlation and the neuron-overlap gradient survive; if a language with few non-shared WALS features behaves like a distant language, or if the correlation vanishes under an alternative distance measure (e.g., human-rated typological similarity), the paper's central claim would be refuted. A second direct falsifier: replace NLLB primes with translations that are known to preserve vs. distort the target English structure; if dis","tokens_in":13348,"feed_emoji":"🌐","tokens_out":1900,"duration_ms":25968,"temperature":0.7,"pith_summary":"This paper tries to establish that crosslinguistic influence (CLI) in language models is not a haphazard side effect of limited capacity, but a structured phenomenon driven by two factors: how dominant the first language is and how proficient the second language is. Using bilingual models trained on one language first and a second language later, the authors show that the more entrenched the L1, the stronger the correlation between structural interference and typological distance between L1 and L2. They also show that a late-introduced L2 has less capacity for positive transfer, making it vulnerable to persistent interference from distant L1s. The paper further provides mechanistic evidence: L1 is co-activated during L2 processing, and syntactically similar L1s share more L2 neurons. If correct, this means training order and language similarity jointly determine how much a model's second language is shaped by its first, giving researchers a controlled way to study bilingualism.","feed_headline":"A dominant first language reshapes a model's second language grammar","feed_subtitle":"Bilingual language models show predictable interference that scales with syntactic distance—a controlled window into human bilingualism.","key_machinery":"The central object is the Step of Exposure (SoE)—the training step at which the L2 is introduced—which jointly manipulates L1 dominance and L2 proficiency. The argument runs on several interlocking tools: syntactic distance measured as the number of non-shared WALS features between L1 and English; crosslinguistic structural priming using NLLB translations of BLiMP sentences as primes; LogitLens to decode L1 token presence during L2 processing; and Parallel Language-specific Neuron Detection (PLND) to identify L2 neurons and measure their overlap across L1 models. These tools connect behavioral accuracy differences (CLI scores) to internal architectural signatures (L1 co-activation, neuron ov","core_discovery":"The central claim is that CLI in language models follows a predictable gradient: as the L1 becomes more dominant (operationalized as later introduction of the L2 during training), the correlation between structural transfer and syntactic distance strengthens dramatically, while reduced L2 proficiency suppresses the model's ability to benefit from positive transfer. Crosslinguistic structural priming—prepending an L1 translation of a grammatical sentence—amplifies these effects, facilitating processing for similar languages and degrading it for distant ones. Mechanistically, the paper finds that the L1 is physically co-activated in hidden states during L2 processing (measured via LogitLens) a","pith_inferences":["An extension the paper leaves implicit: if CLI scales with syntactic distance and dominance, then multilingual models trained with a dominant pivot language (e.g., English-first) may systematically exhibit predictable negative transfer for structurally distant languages—a testable prediction for large-scale multilingual pretraining.","One could test the causal claim more directly by independently manipulating L2 proficiency (e.g., by varying L2 training volume while holding SoE constant) to see whether the distance gradient persists when dominance is not confounded with proficiency.","The WALS-feature distance measure is a coarse proxy; a stronger test would use a behavioral or human-judgment-based syntactic similarity measure, or a larger set of L1s, to check whether the correlation between distance and interference is robust across alternative distance metrics.","If the mechanistic account is right, then targeted interventions—such as pruning or dampening L1-specific neurons during L2 inference—should predictably modulate CLI, offering a direct causal test of the neuron-overlap claim."],"forward_implications":["If L1 entrenchment amplifies distance-correlated interference, then the order and timing of languages in multilingual pretraining is not neutral: a curriculum that entrenches one language first will systematically shape downstream performance in the second.","Crosslinguistic priming effects that scale with syntactic distance give a behavioral test for whether a model has internalized shared syntactic representations, making priming a usable diagnostic for transfer in multilingual models.","The neuron-overlap result implies that typological similarity physically determines how much neural circuitry is shared between languages, which could inform architectural choices for low-resource language transfer.","The romanization finding indicates that orthographic overlap facilitates CLI, suggesting that surface script similarity is a controllable factor in cross-lingual transfer, not just deep syntax.","The asymmetry result—shared structures prime bidirectionally but ungrammatical structures only from dominant L1 to L2—suggests that the shared-versus-connected debate about bilingual syntax can be resolved by considering structural overlap rather than a single universal account."],"fun_headline_variants":["How a dominant first language warps a model's second language","Bilingual AI: L1 dominance predicts grammar interference","Syntactic distance drives crosslinguistic interference in models","L1 dominance and proficiency shape L2 grammar in neural nets","In bilingual models, L1's gravity bends L2 grammar"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire distance-gradient analysis rests on assuming that the number of non-shared WALS features between L1 and L2 is a faithful measure of the syntactic distance that a language model's learning is actually sensitive to; if that feature-based distance does not capture what the model's representations encode, the central correlations and the neuron-overlap interpretation lose their meaning.","fun_headline_variants_meta":{"raw":{"variants":["How a dominant first language warps a model's second language","Bilingual AI: L1 dominance predicts grammar interference","Syntactic distance drives crosslinguistic interference in models","L1 dominance and proficiency shape L2 grammar in neural nets","In bilingual models, L1's gravity bends L2 grammar"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1173,"prompt_tokens":797,"completion_tokens":376,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":293}},"tokens_in":541,"tokens_out":376,"duration_ms":4882,"temperature":1.0,"reasoning_tokens":293,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T06:53:11.437327+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same bilingual framework on a substantially larger set of L1s (e.g., 20 languages with densely sampled syntactic distances) and check whether the CLI-distance correlation and the neuron-overlap gradient survive; if a language with few non-shared WALS features behaves like a distant language, or if the correlation vanishes under an alternative distance measure (e.g., human-rated typological similarity), the paper's central claim would be refuted. A second direct falsifier: replace NLLB primes with translations that are known to preserve vs. distort the target English structure; if dis","supporting_citations":[],"review_version":1}