{"id":"d612ff22-7fa9-4f4d-961e-246c054b3e04","arxiv_id":"2501.14491","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Cross-lingual transfer success is best predicted by syntactic similarity for POS tagging and parsing, by trigram overlap for n-gram topic models, and by mBERT pretraining coverage for mBERT-based topic models.","lead":"This paper tests how well different measures of linguistic similarity predict cross-lingual transfer for 263 languages across POS tagging, dependency parsing, and topic classification. It finds that the best predictor depends on the task and on whether the model uses monolingual or multilingual input representations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missingness in lang2vec syntactic features may bias the finding that syntactic similarity best predicts parsing, and the mixed-effects controls use a tiny non-representative subset; robustness to imputation or high-coverage subsets is untested.","rationale":"Good-faith reading: the paper is an empirical study with a modest, falsifiable claim, and it includes multiple internal checks (cross-model agreement with de Vries et al., the topics-mbert control, per-language correlation tables, and mixed-effects models). I considered whether the task-architecture confound (UDPipe vs MLP) is more load-bearing, but the de Vries comparison and the topics-mbert control partially mitigate it, and the paper carefully avoids causal language. The missingness concern is stronger because it targets the very measure (syn) whose ranking distinguishes the grammatical tasks from topic classification, and because the paper's own data show the effect is much weaker in the non-mBERT subset—exactly where missingness is likely highest. The reader identified the same data-quality premise, so my verdict is unchanged: CONDITIONAL. The proposed test (high-coverage subset + imputed full-set mixed model) would settle whether the concern lands or is resolved.","tokens_in":49075,"tokens_out":6125,"duration_ms":62181,"concrete_test":"Recompute r_avg_syn for LAS and POS on the subset of language pairs with ≥80% syntactic feature coverage (the current average is 63%) and re-fit the §E.2 mixed-effects model on an imputed featurization (e.g., using data2lang2vec, cited as concurrent in the paper) so that all 70×153 UD language pairs are retained. If syn remains the top-ranked predictor in both the high-coverage subset and the imputed full-set model, the concern is resolved; if syn drops below gb or gen, or if the ranking flips for non-mBERT test languages, the central claim must be qualified by resource/documentation level.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that syntactic similarity (syn) is the most predictive measure for POS tagging and dependency parsing rests on correlations computed over language pairs with at least 50% feature coverage (§3.3), yet lang2vec syn coverage averages only 63%, and missingness is almost certainly non-random: typological descriptions are fuller for well-documented, higher-resource languages. The paper's own supplementary results (Table 13) show r_avg_syn for LAS is 0.36 for test languages outside mBERT's pretraining data versus 0.69 for mBERT-covered languages, so the headline effect is concentrated in the resource-rich subset. The only multivariate analysis that can separate syn from correlated measures (phylogenetic relatedness, lexical similarity, mBERT coverage) is the mixed-effects model in §E.2, but it is fit on just 19 training and 31 test languages because of missing data. If missingness correlates with language family, script, or resource level, the measured ranking in Figure 3 may not generalize to low-resource target languages, which is precisely the practical setting the paper addresses. Because syn is the measure that differentiates the grammatical tasks from topic classification, a bias here directly threatens the task-dependence claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies zero-shot cross-lingual transfer across 263 languages from 33 families for three tasks: POS tagging, dependency parsing, and topic classification. It computes several linguistic and dataset-based similarity measures (Grambank, lang2vec syntactic/phonological/phonetic, lexical, phylogenetic, geographic, character/word/trigram overlap) and correlates them with transfer performance at the test-language level. The central claim is that the most predictive similarity measure is not universal: syntactic similarity is most predictive for POS tagging and dependency parsing, trigram overlap is most predictive for n-gram-based topic classification, and mBERT pretraining coverage dominates for mBERT-based topic classification. The paper also derives practical implications for source-language selection. The analyses include per-test-language correlations with confidence intervals (Appendix E.3), mixed-effects models (Appendix E.2), and comparisons across tasks and input representations.","tokens_in":49329,"tokens_out":3150,"duration_ms":31540,"significance":"If the findings hold, they have practical value for source-language selection in multilingual NLP, and they clarify contradictions in prior work by showing that task and input representation condition which similarity measure is the best predictor. The study's scale (263 languages, three tasks, multiple similarity measures, explicit handling of correlations between measures) is a strong point, and the paper is transparent about its data sources, exclusions, and limitations. The main descriptive result is supported by a substantial empirical apparatus, including per-test-language correlations and confidence intervals, which is more rigorous than much prior work in this area. The paper also makes a useful methodological contribution by comparing dataset-dependent and dataset-independent similarity measures in a unified framework.","major_comments":[{"comment":"The central claim that syntactic similarity (syn) is the most predictive measure for POS tagging and dependency parsing is load-bearing for the paper's task-dependence conclusion, but it may be driven by the resource-rich subset of languages. Appendix E.3 (Table 13) shows that the mean LAS correlation with syn is 0.69 for test languages in mBERT's pretraining data but only 0.36 for test languages outside it; for POS the corresponding figures are 0.46 and 0.20. Because syn coverage in lang2vec averages only 63% and missingness is likely correlated with language documentation/resource level, the per-test-language correlations for low-resource targets are not only weaker but may be non-representative. The paper should either provide a robustness analysis using imputed features or a high-coverage subset of languages, or explicitly restrict the claim to languages covered by mBERT and discuss the implications for low-resource target languages. Without this, the ranking in Figure 3 may not generalize to the practical setting the paper addresses.","section":"§3.3, Table 13, Figure 3"},{"comment":"The mixed-effects model is the only multivariate analysis that separates syn from correlated predictors such as phylogenetic relatedness, lexical similarity, and mBERT coverage, but it is fit on only 19 training and 31 test languages for the grammatical tasks (42 for topic classification) because of missing data. This is a small, non-representative subset, and the model still exhibits substantial collinearity between gen and lex and between gb and syn (correlations between -0.769 and -0.819 for gen-lex and between -0.442 and -0.509 for gb-syn). I acknowledge the paper's statement that significance values are based on model comparison, which is a reasonable approach, but the generalization of the multivariate conclusions to the full language set is questionable. Adding a sensitivity analysis on a more complete (even if smaller) set of languages with complete feature coverage, or a careful discussion of how missingness could affect the mixed-effects estimates, would substantially strengthen the paper.","section":"§E.2, Table 11"}],"minor_comments":[{"comment":"In the sentence 'If we consider our own POS tagging results but only select the subset of languages that was used in de de Vries et al.'s experiments', there is a duplicated 'de'; it should read 'de Vries et al.'s experiments'.","section":"§4.1"},{"comment":"The label 'word*' in Figure 3 is defined only in the caption as 'overlap between words (wor; UD tasks), trigrams (tri; topics-base/translit), and subword tokens (swt; topics-mbert)'. Consider adding the mapping directly in the figure or a more explicit legend, since the current notation may confuse readers.","section":"§5.1, Figure 3"},{"comment":"The source-selection analysis would benefit from a random-source baseline. Without knowing the average performance loss when picking a source language uniformly at random, it is difficult to calibrate whether a given loss is 'small' or 'adequate'. The relative ordering of measures is clear, but the practical claim about absolute adequacy is not fully quantified.","section":"§5.2.1, Table 3"},{"comment":"The paper treats correlation coefficients with p-values of at least 0.05 as zero. This is a conservative choice, but reporting the raw p-values or providing a significance asterisk in the main figures would make the analysis more transparent, especially for readers who want to assess borderline correlations.","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid and well-executed empirical study with a clear central research question and appropriate methodological care. My main concern is the potential bias from missing data in the typological databases, which may affect the headline claim about syntactic similarity. I would encourage the authors to add a robustness analysis or substantially temper the claim. The paper also fits the journal's scope well, and I do not see any issue with novelty or attribution. The minor issues listed are presentation-level and should be easy to address."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a serious referee. The headline result—that the best similarity predictor for zero-shot transfer depends on the task and the input representation—is a real empirical contribution, backed by a broader language and task coverage than anything comparable in the literature. I trust the main descriptive claim.\n\nThe paper does several things well. The per-test-language correlations with confidence intervals in Appendix E.3 are the right way to avoid aggregation artifacts, and the mixed-effects models, though small, are an honest attempt to separate correlated predictors. The practical heuristics in Section 5.2 are useful, and the failure modes (e.g., high character overlap not guaranteeing good transfer) are clearly shown.\n\nThe stress-test concern about missingness in lang2vec syn features is legitimate but does not sink the paper. The authors report the LAS correlation with syn separately for mBERT-covered (0.69) and non-covered (0.36) test languages; syn is still the strongest predictor in the non-covered subset, so the task-dependence claim survives, though the effect is weaker exactly in the low-resource setting practitioners care about. The mixed-effects model is fit on only 19 training and 31 test languages due to missing data—that is a real limitation and should be reported more prominently in the main text. One model per task and the absence of a random-source baseline for the heuristics are minor gaps. No code release is mildly annoying for a paper of this scope.\n\nRecommendation: conditional accept. The paper deserves referee time and will be useful to anyone working on source-language selection. I would ask the authors to add robustness checks for missingness (e.g., imputation or high-coverage subsets) and to release code.","headline":"Solid empirical paper; the task-dependence claim survives the missingness concern, though the effect weakens for low-resource targets.","tokens_in":49824,"tokens_out":1978,"would_cite":true,"duration_ms":18982,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across 263 languages, the best predictor of zero-shot transfer is task-dependent: syntax for parsing, trigram overlap for n-gram topic models, mBERT coverage for mBERT models.","keywords":["cross-lingual transfer","linguistic similarity","source language selection","dependency parsing","POS tagging","topic classification","zero-shot transfer","mBERT pretraining coverage"],"falsifier":"A decisive check is to rerun the exact source-selection heuristic on a held-out set of language families with complete typological feature coverage and to see whether the winner remains syntactic similarity for parsing and trigram overlap for n-gram topic classification; if the ranking flips, the task-dependence claim is an artifact of the included languages.","tokens_in":48896,"feed_emoji":"🌐","tokens_out":11000,"duration_ms":93976,"temperature":0.7,"pith_summary":"The paper asks a practical question: when building a zero-shot NLP system for a low-resource language, how should you choose the language to train on? It answers with a large-scale measurement: across 263 languages and three tasks, the similarity measure that predicts transfer performance depends on the task and on how the input is represented. Syntactic similarity is the strongest guide for POS tagging and dependency parsing; string and lexical similarity lead for n-gram-based topic classification; and for topic classification with multilingual transformer embeddings, no similarity measure is strong, and what matters most is whether the target language was in the model's pretraining data. If these findings hold, source-language selection should be task-aware rather than based on a single all-purpose language distance. The paper also shows that when no task-specific results exist, borrowing the best source from a conceptually similar experiment is a workable fallback.","feed_headline":"Language similarity's best predictor shifts with task and model","feed_subtitle":"Across 263 languages, the best similarity predictor depends on task and input representation, so source selection should be task-aware.","key_machinery":"The load-bearing device is a paired comparison between two matrices: a matrix of zero-shot transfer scores for every source-target language pair, and a series of pairwise similarity matrices computed for the same language pairs. Similarity is quantified with Gower's coefficient, a feature-wise agreement score that ignores missing values, applied to typological feature vectors, plus lexical and phylogenetic distances, geographic distance, and dataset-level character, word, trigram, and subword overlap. The analysis ranks similarity measures by their per-test-language Pearson correlation with transfer scores, then reruns the ranking in a source-selection simulation where the most similar language per measure is chosen as the training language and the loss against the best possible source is measured. That simulation turns the correlation result into an actionable heuristic.","core_discovery":"The paper's central discovery is that the ranking of linguistic-similarity measures as predictors of zero-shot transfer changes with the task and the input representation. For dependency parsing, the strongest predictor is syntactic similarity, with an average Pearson correlation of 0.57 against labeled attachment score, followed by typological-feature similarity at 0.42; for POS tagging the same syntactic measure leads at 0.37. For n-gram-based topic classification, character trigram overlap is the strongest predictor in the original-script setup (0.65), with lexical similarity close behind, while in the transliterated setup lexical similarity leads (0.61); for topic classification built on multilingual transformer embeddings, no similarity measure exceeds 0.27 and performance instead tracks whether source and target languages appear in the model's pretraining data (68.3% average accuracy when both do, 33.2% when neither does). The practical corollary, demonstrated by a source-selection simulation, is that choosing a source by the task-appropriate measure yields small performance losses, and choosing by the results of a conceptually similar experiment is a safe fallback.","pith_inferences":["Beyond the paper, the result implies that a single all-purpose language similarity score cannot serve every NLP task; practical tooling should expose separate similarity layers so users can match the measure to their task and model family.","The winning role of pretraining coverage for transformer-based topic classification suggests that a tokenizer-exposure index, computed directly from the model's vocabulary, might predict transfer for other frozen multilingual encoders and could be tested on the paper's own 194-language matrix.","A natural extension the authors do not run is to check whether the winning measure changes with model scale or architecture, which would show whether the task-dependence they find is tied to mBERT-style tokenization or is a general property of cross-lingual transfer."],"forward_implications":["For POS tagging and dependency parsing, practitioners should rank candidate source languages by syntactic similarity; doing so produced near-best transfer performance among the heuristics tested.","For n-gram-based topic classification, ranking sources by character trigram overlap or lexical similarity is the better heuristic, while for transformer-embedding topic classification the source and target should be checked against the model's pretraining coverage.","If no task-specific experiment is available, using transfer results from a conceptually similar task with similar input representations yields small losses, but using mismatched representations is a poor guide.","Transfer patterns for the two grammatical tasks are highly correlated with each other, so conclusions from one grammatical task are likely to carry over to the other, while the boundary between grammatical and topic-classification tasks matters.","Training dataset size and phonological or phonetic similarity were weak predictors across experiments, so they can be deprioritized in source selection for these tasks."],"supporting_citations":[{"why":"Supplies the syntactic, phonological, and phoneme-inventory feature vectors whose syntactic dimension is the strongest predictor for parsing and POS tagging.","marker":"(Littell et al., 2017)"},{"why":"Provides Grambank's grammatical features, the second-strongest similarity predictor for parsing and POS tagging.","marker":"(Skirgård et al., 2023)"},{"why":"Provides the UDPipe 2 tagger-parser models whose POS and dependency scores form the grammatical transfer matrices.","marker":"(Straka, 2023)"},{"why":"Supplies the SIB-200 parallel topic-classification dataset and the MLP baselines the topic models are compared with.","marker":"(Adelani et al., 2024)"},{"why":"Supplies the mBERT embeddings used by UDPipe 2 and by the mBERT-based topic classifier; pretraining coverage drives that setup's transfer.","marker":"(Devlin et al., 2019)"},{"why":"Gives the similarity coefficient used to convert typological feature vectors into pairwise language scores.","marker":"(Gower, 1971)"},{"why":"Provides the lexical dissimilarity scores underlying the lexical similarity predictor, which leads for transliterated n-gram topic classification.","marker":"(Jäger, 2018a,b)"},{"why":"Supplies the Universal Dependencies 2.14 test splits used for the 153-language POS tagging and dependency parsing experiments.","marker":"(Zeman et al., 2024)"},{"why":"Supplies the prior 65-by-105 POS-transfer matrix whose results correlate strongly with the paper's own, anchoring the comparison across model choices.","marker":"(de Vries et al., 2022)"}],"fun_headline_variants":["Task and setup dictate best similarity predictor for transfer","Which language similarity helps transfer? It depends on task","263 languages: similarity's role in transfer is task-specific","No universal similarity measure for cross-lingual transfer","For cross-lingual transfer, task beats language similarity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rankings assume that languages missing from typological databases are missing at random; if entire families or scripts are underrepresented in those databases, the measured correlations between similarity and transfer could be biased.","fun_headline_variants_meta":{"raw":{"variants":["Task and setup dictate best similarity predictor for transfer","Which language similarity helps transfer? It depends on task","263 languages: similarity's role in transfer is task-specific","No universal similarity measure for cross-lingual transfer","For cross-lingual transfer, task beats language similarity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1354,"prompt_tokens":916,"completion_tokens":438,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":361}},"tokens_in":532,"tokens_out":438,"duration_ms":5164,"temperature":1.0,"reasoning_tokens":361,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:05:29.888311+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check is to rerun the exact source-selection heuristic on a held-out set of language families with complete typological feature coverage and to see whether the winner remains syntactic similarity for parsing and trigram overlap for n-gram topic classification; if the ranking flips, the task-dependence claim is an artifact of the included languages.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the UDPipe 2 tagger-parser models whose POS and dependency scores form the grammatical transfer matrices."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Universal Dependencies 2.14 test splits used for the 153-language POS tagging and dependency parsing experiments."}],"review_version":1}