{"id":"bd162bc3-c5bc-44b0-a012-691913acd0d8","arxiv_id":"2505.13908","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper finds, through XLM-R experiments on WikiANN, that within-family and morphologically similar language pairs transfer better, but it mostly restates prior findings.","lead":"This preprint reports an XLM-R study of zero-shot POS tagging across 15 languages and claims that language family and morphology predict transfer success. It is a rough empirical confirmation of a known result, but the central figures are missing and the data table shows implausible same-language scores.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The only numerical evidence is an internally inconsistent transfer matrix: task is POS in §2.1 but NER in §2.5, the Chinese diagonal is 0.12, and German→Arabic exceeds German→French, contradicting the intra-family claim.","rationale":"The reader's weakest assumption correctly identifies protocol validity as the load-bearing condition, and the paper's own table provides decisive checks that this condition is not met. Section 2.1 commits the experiment to POS tagging on WikiANN, but WikiANN is an NER corpus and Section 2.5 explicitly mentions NER, so the task label is internally inconsistent. The Chinese diagonal value of 0.12 is not a plausible supervised same-language F1 for any standard POS or NER benchmark with XLM-R, and it is markedly lower than every other diagonal value. Finally, the paper's central intra-family claim is contradicted by its own numbers: German→Arabic (0.686) is higher than German→French (0.65), even though Arabic is Afro-Asiatic and French is Indo-European. Since the data table is the only quantitative evidence and the five figures are not present, these inconsistencies are load-bearing. A small reproduction of three cells and one diagonal would settle the issue; until then the claim should not be accepted. The paper does cite existing literature reporting proximity-correlated transfer, so the qualitative direction is not novel, but the present evidence is unusable as reported.","tokens_in":6347,"tokens_out":5636,"duration_ms":55649,"concrete_test":"Reproduce a minimal decisive subset of the table: fine-tune XLM-R-base on German and evaluate on German, French, and Arabic with three seeds, following the stated zero-shot protocol. If the claimed task is POS, WikiANN cannot supply POS labels, so the task mismatch is confirmed; if the task is NER, run WikiANN NER instead. The table is only credible if the German self-score is high (roughly 0.7 or above) and German→French is at least as large as German→Arabic. If the Chinese diagonal also fails to reproduce at a plausible supervised level, the central numerical evidence is invalidated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('typological proximity, rather than raw training data size alone, remains the dominant driver') rests entirely on the pairwise table at the end of §3. For that table to support the claim, it must be a correctly run zero-shot POS-transfer matrix from WikiANN. At least three features show it is not. (1) Task mismatch: §2.1 states POS tagging, but WikiANN is a named-entity dataset and §2.5 explicitly refers to 'cross-lingual NER performance'; POS tagging and NER have different label spaces and different morphological sensitivity. (2) Impossible diagonal: same-language cells should be supervised in-language performance and should be high, since the model is fine-tuned on that language and evaluated on its own test set; the Chinese diagonal is 0.12, far below every other diagonal and below many cross-lingual cells, and no explanation is offered. (3) Internal contradiction with the headline pattern: German→Arabic is 0.686 while German→French is 0.65 and German→Spanish is 0.72, so a cross-family pair outscores a within-family pair; row/column asymmetries such as French→Arabic 0.649 versus Arabic→French 0.55 are never discussed. With no error bars, no seed variance, and undefined 'Base languages', the table cannot establish that family or morphology dominates transfer. The five figures are also missing, so the narrative effect sizes of §3 cannot be checked independently.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies zero-shot cross-lingual transfer with XLM-R across 15 languages from several families, using WikiANN as the data source. The procedure is to fine-tune XLM-R on a source language and evaluate on target languages, reporting a pairwise transfer matrix of average F1 values. The authors claim that intra-family transfer outperforms cross-family transfer, that morphological properties such as fusional or agglutinative structure and gender marking influence transfer, and that linguistic distance correlates negatively with transfer success. The central conclusion, stated in Section 4.1, is that typological proximity, rather than raw training data size alone, remains the dominant driver of effective knowledge transfer between languages.","tokens_in":6623,"tokens_out":4298,"duration_ms":39483,"significance":"If the empirical results were reliable, the paper would add supporting evidence for a claim that is already established in the multilingual NLP literature, for example by Lauscher et al. (2020) and Pires et al. (2019). The paper does not propose a new method, dataset, or theoretical framework, so its significance rests entirely on the credibility of its experiments. A strength is that the authors include their full transfer matrix and state that each experiment was run three times, which makes the empirical basis transparent. However, the reported evidence is internally inconsistent and partly missing, so the central claim is not currently supported. The paper also has no machine-checked proofs or released code; its contribution is purely empirical.","major_comments":[{"comment":"Section 2.1 defines the task as Part-of-Speech (POS) tagging and names WikiANN as the dataset, but WikiANN is a named-entity recognition dataset and Section 2.5 explicitly refers to 'cross-lingual NER performance.' POS tagging and NER have different label spaces and different sensitivity to morphology, so the reported F1 table cannot be interpreted without knowing which task was actually run. This inconsistency undermines the central comparison because every conclusion in Section 3 depends on the transfer matrix.","section":"§2.1 and §2.5"},{"comment":"The diagonal entry for Chinese (source Chinese, target Chinese) is 0.12, far below every other diagonal value and below many cross-lingual cells; no explanation is offered. Since same-language evaluation should reflect supervised fine-tuning performance, this value is implausible and suggests a data-processing or labeling error. Additionally, the table has rows for Tamil, Korean, and Japanese but no columns for these languages, so it is not a complete pairwise matrix and the missing entries are not discussed.","section":"Data table (unnumbered, after References)"},{"comment":"Figures 1–5 appear only as placeholder captions; no actual plots are included. The narrative in Sections 3.1–3.4 reports specific effect sizes, including intra-family scores surpassing 0.70, cross-family drops exceeding 0.15, and a negative distance–transfer correlation, none of which can be verified without the figures. Because these figures are the only quantitative support for the claims in Section 4.1, the central conclusion is not supported by the submitted manuscript.","section":"§3 and §4.1"},{"comment":"The table contains direct evidence against the 'intra-family dominance' claim: German→Arabic (0.686) is a cross-family pair that outscores German→French (0.65), a within-family pair; similarly, French→Arabic (0.649) is much higher than Arabic→French (0.55). These asymmetries are not discussed. Section 2.4 states that each experiment was run three times, but no error bars, standard deviations, or seed-level results are reported, so the table cannot establish that family or morphology is the dominant factor.","section":"Data table (unnumbered, after References)"}],"minor_comments":[{"comment":"The column header 'Finish' should read 'Finnish'; this is a factual error that should be corrected.","section":"Data table header"},{"comment":"The caption says the table reports 'average F1 Accuracy and Precision,' but it is unclear whether the numbers are F1 scores, accuracy, precision, or some combination; the metric should be stated precisely.","section":"Data table caption"},{"comment":"The text labels English as 'Indo-Aryan'; English is Indo-European, and English does not appear in the reported table, so the example pair 'English → German' is misleading.","section":"§2.1"},{"comment":"The text refers to 'geographical (or typological) distance' but never defines the distance metric used in Figure 5; geography and typology are distinct notions and should not be conflated without explanation.","section":"§3.3"},{"comment":"There are numerous typos and grammatical errors, including 'Base languges,' 'out primary reason,' 'we analyzed on15 languages,' and 'compare it to very other language'; the manuscript would benefit from a careful proofreading pass.","section":"Throughout"},{"comment":"The figures are referenced out of order in the text (Figure 1, Figure 3, Figure 2, Figure 4, Figure 5), which makes the narrative harder to follow and should be renumbered or reordered.","section":"§3"}],"recommendation":"reject","confidential_remarks":"The empirical core of this manuscript is not in a publishable state: the figures are missing, the task description is inconsistent between POS tagging and NER, and the transfer matrix contains implausible diagonal values and internal contradictions. If the authors can rerun the experiments, resolve the task ambiguity, and provide the actual figures and seed-level statistics, a future resubmission might be considered, but the current submission does not meet the bar for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What should you know about arXiv:2505.13908? It's an under-specified empirical paper claiming that typological and family proximity drive zero-shot transfer in XLM-R. The claim itself is not new—their own references [11] and [14] state it—and the only new evidence, a pairwise transfer table, is too inconsistent to support anything.\n\nThe paper does some things right. The authors have read the right literature: Lauscher, Pires, Ponti, Üstün are all there and cited appropriately. Organizing languages by family and morphology is a sensible frame, and measuring intra- vs cross-family transfer with a fixed model is reasonable. If the table were clean, it would be a small but usable data point.\n\nIt is not clean. Section 2.1 says the task is POS tagging, but WikiANN is an NER dataset and Section 2.5 explicitly says \"cross-lingual NER performance.\" The same-language diagonal for Chinese is 0.12, which is impossible for a model fine-tuned and evaluated on Chinese; every other diagonal is above 0.6. German→Arabic (0.686) outscores German→French (0.65), directly contradicting the paper's intra-family claim. The table has no error bars, even though §2.4 says each experiment was run three times. \"Base languages\" are never defined. The example pair in §2.1 calls English \"Indo-Aryan\"—it's Indo-European but not Indo-Aryan. \"Finish\" is used for Finnish, and the table has 17 languages while §2.5 says 15. Figures 1–5 are placeholder captions with no images, so the effect sizes described in §3 cannot be checked.\n\nThat list is not a pile-on. Any one of these might be a typo. Together, they mean the central comparison rests on evidence that is internally contradictory. The paper's conclusion—typological proximity dominates transfer—is almost certainly true in the literature, but this manuscript does not add a reliable measurement of it.\n\nThe citation pattern is conventional and not self-serving. The writing is rough but readable. The problem is that the empirical core is unusable as written.\n\nWho is this for? Someone teaching a course on multilingual NLP might assign it as an example of how not to report transfer experiments. A researcher would not get a new result from it. I would not send it to peer review in this state; the authors should fix the task description, produce the figures, release code and per-seed results, and reconcile the table with the stated protocol. If they do that, a corrected version could be a minor empirical contribution. As it stands, it's a reject.","headline":"The central claim repeats the paper's own references, and the only new evidence—a transfer table with an impossible Chinese diagonal, a POS/NER task mismatch, and a cross-family pair outsoring within-family pairs—cannot support it.","tokens_in":7140,"tokens_out":3084,"would_cite":false,"duration_ms":28719,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that typological proximity—shared language family and shared morphology—rather than raw pretraining-data size is the dominant driver of zero-shot cross-lingual transfer in XLM-R, based on fine-tuning experiments over 15…","keywords":["cross-lingual transfer","multilingual NLP","language families","morphology","zero-shot transfer","XLM-R","part-of-speech tagging","typological distance"],"falsifier":"Recompute the same 15-language zero-shot transfer matrix under a single, unambiguous task definition with the source languages and train/test splits specified; if the Chinese diagonal is not near 0.12, or if intra-family advantages shrink once the protocol is fixed, the central comparison collapses. A sharper experiment would hold training data constant and compare genealogically close but typologically distant pairs against genealogically distant but morphologically close pairs (for example Turkish, Finnish, and Tamil, all agglutinative); the paper's claim predicts the morphologically close pairs transfer better regardless of family.","tokens_in":6148,"feed_emoji":"🌍","tokens_out":9464,"duration_ms":86529,"temperature":0.7,"pith_summary":"The paper sets out to show that when a massively multilingual model transfers knowledge to a language it never saw during fine-tuning, success is governed more by how similar the languages are than by how much pretraining data the model had. The authors fine-tune XLM-R on one source language at a time and evaluate zero-shot on the other 14 languages in a 15-language set, using WikiANN annotations for part-of-speech tagging. They report that intra-family source–target pairs usually exceed 0.70 F1 while cross-family pairs often fall below 0.50, and that fusional and agglutinative languages transfer better than isolating ones such as Chinese. The conclusion is that typological proximity, rather than raw training-data size alone, remains the dominant driver of effective knowledge transfer between languages. If true, this gives NLP practitioners a concrete rule for choosing source languages and a reason to build morphology-aware training and tokenization.","feed_headline":"Language family, not data size, drives cross-language transfer","feed_subtitle":"Fine-tuning XLM-R on a sibling language lifts POS-tagging F1 past 0.70; distant pairs fall below 0.50.","key_machinery":"The central object is XLM-R, a 12-layer multilingual Transformer pretrained on about 100 languages with CommonCrawl data; it is the model being fine-tuned, and its learned subword representations are what the paper claims get reused. The argument runs through a transfer matrix built by fine-tuning on each source language's WikiANN training set and evaluating on each target language's test set with no target training sentences. The paper interprets that matrix with three categorical lenses—language family, morphological type (fusional, agglutinative, isolating), and presence of gender marking—plus a distance axis; together these lenses turn raw F1 numbers into evidence that structural affinity, not data volume, drives transfer.","core_discovery":"On the authors' own terms, the discovery is that XLM-R's zero-shot transfer succeeds to the extent that source and target languages are typologically close: language-family membership and shared morphological features predict F1 more strongly than corpus volume. The evidence is a 15×15 transfer matrix in which intra-family scores, such as Spanish→French, typically pass 0.70, whereas cross-family pairs such as Arabic→Japanese can drop below 0.50. Fusional and agglutinative languages cluster near 0.60–0.65 in intra-family transfer, isolating languages like Chinese near 0.40, and gendered languages beat non-gendered ones by 0.05–0.10 within a family. The paper reads the matrix as showing that shared inflectional systems, affixation patterns, and gender marking enable parameter reuse, while tonal and isolating structure leaves the model without a bridge, so large pretraining corpora cannot erase structural divergence.","pith_inferences":["The paper does not spell out an automated source-selection rule, but its claim implies one: for any low-resource target, pick the source language with the smallest morphological distance rather than the largest corpus, and test this by swapping sources while holding the target fixed.","The paper's distance plot mixes genealogical and morphological distance; separating these with strict typological features would show whether shared agglutinative structure alone (Turkish, Finnish, Tamil) drives transfer without any family tie, which the outlier pattern hints at.","If the claim generalizes beyond these 15 languages, zero-shot evaluations should be reported as family-by-family submatrices; aggregate scores will conceal which languages are being left behind, and model rankings may flip depending on the family mix of the benchmark."],"forward_implications":["Choosing a source language from the target's own family should improve zero-shot performance; the paper explicitly suggests using Finnish rather than English when building a model for Estonian.","Morphological similarity can partly compensate for genealogical distance, so Turkish and Hungarian—both agglutinative—should transfer better to each other than their family separation alone would predict.","Isolating and tonal languages such as Chinese are the most difficult zero-shot endpoints, so they are the languages most likely to need family-specific models or morphology-aware tokenization.","Injecting morphological information—through analyzers, character-level modeling, or multi-task morphological tagging—should narrow the transfer gap for morphologically complex low-resource languages.","Zero-shot benchmarks should be stratified by family and morphological profile, since overall averages can hide systematic failure on typologically distant languages."],"supporting_citations":[{"why":"Supplies the empirical finding that zero-shot transfer success correlates with linguistic proximity between source and target languages, which the paper extends to families and morphology.","marker":"[11]"},{"why":"Provides the premise that transfer works best when languages are similar in structure or lineage, the idea the paper tests on XLM-R.","marker":"[14]"},{"why":"Defines XLM-R, the 12-layer multilingual Transformer whose zero-shot behavior the paper measures.","marker":"[3]"},{"why":"Supplies typological, geographical, and phylogenetic language vectors, the underpinning for the paper's linguistic-distance analysis.","marker":"[13]"}],"fun_headline_variants":["Morphology beats corpus size for cross-language transfer","Language family proximity, not data, predicts zero-shot transfer","Shared morphology boosts XLM-R transfer; distant pairs flop","Typological closeness drives cross-lingual transfer, not corpus volume"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every number in the reported table is a correctly computed zero-shot F1 score for one consistent task, and the paper itself weakens that premise by giving Chinese a diagonal of 0.12, by calling the task POS tagging in one section and NER in another, and by leaving the source 'Base languages' undefined.","fun_headline_variants_meta":{"raw":{"variants":["Morphology beats corpus size for cross-language transfer","Language family proximity, not data, predicts zero-shot transfer","Shared morphology boosts XLM-R transfer; distant pairs flop","Typological closeness drives cross-lingual transfer, not corpus volume"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1173,"prompt_tokens":878,"completion_tokens":295,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":226}},"tokens_in":494,"tokens_out":295,"duration_ms":3116,"temperature":1.0,"reasoning_tokens":226,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:07:01.313874+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the same 15-language zero-shot transfer matrix under a single, unambiguous task definition with the source languages and train/test splits specified; if the Chinese diagonal is not near 0.12, or if intra-family advantages shrink once the protocol is fixed, the central comparison collapses. A sharper experiment would hold training data constant and compare genealogically close but typologically distant pairs against genealogically distant but morphologically close pairs (for example Turkish, Finnish, and Tamil, all agglutinative); the paper's claim predicts the morphologically close pairs transfer better regardless of family.","supporting_citations":[{"cited_title":"From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual Transformers","cited_arxiv_id":null,"evidence_quote":"Supplies the empirical finding that zero-shot transfer success correlates with linguistic proximity between source and target languages, which the paper extends to families and morphology."},{"cited_title":"URIEL and LANG2VEC: Representing Languages as Typological, Geographical, and Phylogenetic Vectors","cited_arxiv_id":null,"evidence_quote":"Supplies typological, geographical, and phylogenetic language vectors, the underpinning for the paper's linguistic-distance analysis."}],"review_version":1}