{"id":"a54dc325-7a98-4394-84c7-662f2b72b5c8","arxiv_id":"2608.03446","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"English-based cross-lingual alignment predicts LLM translation quality as well as or better than direct source-target alignment, supporting the English-pivot hypothesis.","lead":"This paper compares 27 ways of measuring how closely a multilingual AI model's internal representations of different languages line up with each other, and tests which ones predict real performance on classification and translation. It finds that alignment with English predicts translation quality as well as or better than direct source-to-target alignment, except for closely related languages.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PMI target-language independence is the load-bearing assumption for the English-pivot evidence; it is validated circularly and fails for one model.","rationale":"The reader's judgment is sound. The central claim has two components: a predictive claim (English-side CLA predicts translation quality) and an interpretive claim (English pivot). The predictive claim is supported by multiple correlations and an out-of-domain BOUQuET validation, so I would not reject it. The interpretive claim, however, depends on the Section 6.3 PMI analysis, and that analysis rests on PMI being a near-target-language-independent measure of translation quality. This is the least secure link: the paper's own validation is indirect, fails for one model on the main dataset, and the Limitations disclaim formal independence. The Appendix C 'one-sided implication' is circular because it uses the very mixed-target correlation that is supposed to be explained. A target-language fixed-effect re-analysis would settle whether the cross-target correlation is real or an ecological artifact. I also note secondary concerns (in-sample selection of weights, collinearity of alignment scores) but they do not change the verdict conditional on the PMI test. The released code makes the proposed check feasible, so the conditional verdict remains appropriate.","tokens_in":33035,"tokens_out":8759,"duration_ms":98122,"concrete_test":"Recompute the Section 6.3 correlation and alpha/beta search after partialling out target-language identity: subtract per-target-language means from PMI and from the combined alignment score (or include target-language dummy variables) before computing Pearson correlations and optimal weights. If the correlation drops substantially (e.g., from ~0.9 to below ~0.7) or the optimal src-tgt weight moves away from zero, the English-pivot evidence is a target-language confound rather than a genuine predictive relationship. This test is directly runnable from the released code and data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central English-pivot claim (Section 6.3, Table 7) requires comparing translation quality across target languages, which is only possible through the proposed PMI metric. The validity of PMI as a target-language-independent proxy is the load-bearing assumption, and it is not established. Within-target validation (Table 4) is averaged over targets and has a striking failure: gemma-3-instruct on Flores correlates only 26.2% with chrF, yet this model is still used in the main PMI analysis; the Figure 8 note concedes this leads to 'wrong parameter estimation.' The target-language independence argument in Appendix C is circular: it takes the strong mixed-target correlation with the combined alignment as evidence of independence, but that is the very correlation used to infer the pivot. Limitations explicitly state 'full independence cannot be formally established.' If PMI retains any target-language-specific offset (from tokenization, calibration, length normalization, script), then en-tgt CLA—which varies only across target languages—will correlate with PMI at the language-pair level even when English alignment carries no information about translation quality. The near-zero src-tgt weight would then be a confound, not evidence for an internal English pivot.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a comparative analysis of 27 cross-lingual alignment (CLA) score variants—obtained by combining five sentence-embedding extraction methods (including a newly proposed few-shot method), six alignment metrics, and one tokenizer-level metric—for predicting LLM performance on SIB-200 topic classification, Belebele reading comprehension, and Flores/BOUQuET machine translation. The experiments cover six model variants (Qwen3, gemma-3, Ministral-3, base and instruct) and 44 languages. To analyze translation performance across target languages, the paper introduces a PMI-based translation quality metric intended to be less target-language-dependent than chrF. The central claim is that English-centric CLAs (src-en and en-tgt) predict translation quality comparably to or better than source-target CLA, providing new evidence that LLMs use English as an internal pivot language. The analysis also includes an out-of-domain validation on BOUQuET and a related-languages side experiment.","tokens_in":33288,"tokens_out":6206,"duration_ms":67348,"significance":"The paper is a useful and broad empirical contribution. It substantially extends the study of CLA scores from classification to translation, provides the first comparison of 27 CLA variants on modern decoder-only LLMs, and releases code. The out-of-domain validation on BOUQuET is a notable strength, as is the inclusion of six model variants and 44 languages. The less language-dependent PMI translation metric is an interesting idea, and if its target-language independence were established, the English-pivot result would be a significant addition to the literature. However, the load-bearing evidence for the pivot claim relies critically on the PMI metric's cross-target comparability, and the current validation of that property is not conclusive. The paper also contains a circular evaluation involving the Dali metric and Belebele, and the in-sample selection of the best of 27 variants and ensemble weights is not accounted for. These issues are fixable and do not negate the value of the comparative results, but they currently prevent the strongest conclusions from being accepted as stated.","major_comments":[{"comment":"The claim that en-tgt CLA predicts translation quality—and hence the English-pivot conclusion—depends on the PMI metric being comparable across target languages. The within-target correlation with chrF (Table 4) validates PMI only as a quality proxy for a fixed target; it does not validate cross-target comparability. The mixed-target evidence in Appendix C is circular: the strong correlation of the combined alignment (which includes en-tgt) with PMI is used both to establish PMI's target-language independence and to infer the pivot. The Limitations section explicitly concedes that \"full independence cannot be formally established.\" Moreover, the Flores correlation for gemma-3-instruct is only 26.2% (Table 4), and the note to Figure 8 concedes that this leads to \"wrong parameter estimation.\" Since the cross-target PMI analysis in Table 7 is the primary support for the English-pivot interp","section":"Sections 6.1, 6.3, Appendix C, Limitations"},{"comment":"Dali scores are estimated on the Belebele dataset (Section 3.1: \"Dali scores are estimated on the Belebele dataset, instead of Flores\") and then correlated with Belebele accuracy in Section 5. For example, Table 9 reports Dali correlations as high as 93.1% (Ministral-3-instruct). This is circular: the predictor is derived from the same task and data distribution as the target, so the high correlation does not demonstrate that Dali predicts reading comprehension in general; it may simply reflect the model's retrieval accuracy on the same benchmark. This is especially relevant because Dali is included in the 27 variants and the paper recommends retrieval-based metrics in Section 7. Please recompute Dali on a separate dataset (or held-out portion) and re-evaluate, or clearly exclude the Belebele-based Dali results from the correlational claims.","section":"Sections 3.1 and 5, Table 9"},{"comment":"The headline correlations in Tables 6 and 7 are the maximum over 27 CLA variants and over grid-searched ensemble weights (step size 1/25), all selected on the Flores devtest set. The BOUQuET validation later applied to the selected configuration is a good practice, but the reported in-sample numbers (e.g., ρ⊕ = 91.2% in Table 7) are subject to multiple comparisons; no correction or stability analysis is provided. More importantly, the conclusion that the src-tgt weight is \"consistently near-zero or zero\" (Section 6.3) is based on the argmax weights, yet Table 13 shows that adding src-tgt improves the correlation over src-en alone by at most a few points and often by zero. The weight landscape is likely flat, so the zero-weight result is underdetermined. Please report the distribution of optimal weights, confidence intervals, or a bootstrap analysis to demonstrate that the weights are not","section":"Sections 6.2 and 6.3, Tables 6, 7, 13, 15, 17"},{"comment":"The related-languages experiment (Section 6.4) is presented as a boundary condition on the English-pivot claim: for close language pairs, the src-tgt weight becomes dominant (Table 7 bottom). However, this experiment uses only 14 data points (7 pairs, base models only), and the Limitations admit that \"token overlap between related languages may skew the PMI and CHRF distributions.\" Given that PMI's target-language independence is already under question, the related-languages result cannot distinguish between the interpretation \"no English pivot when translation is trivial\" and the artifact \"PMI is unreliable for high-overlap language pairs.\" This does not invalidate the experiment, but it should be interpreted much more cautiously, or supplemented with an analysis that controls for token overlap.","section":"Section 6.4 and Limitations"}],"minor_comments":[{"comment":"The caption says \"optimal convex combination of en-tgt and src-tgt\" but Equation (6) and the surrounding text use src-en and src-tgt. This is a typo and should be corrected.","section":"Table 6 caption"},{"comment":"The x-axis label appears to be missing the symbol for alpha (shows \"weight for src-en ( )\"). Please fix the formatting.","section":"Figure 4"},{"comment":"The phrasing \"we propose that a strong correlation within mixed target languages ... indicates target-language independence. Note that this is only a one-sided implication\" is logically confusing: if the correlation is the same one used to infer the pivot, the implication is not merely one-sided but circular. The distinction should be clarified.","section":"Appendix C"},{"comment":"The paper repeatedly refers to \"CHRF\" with uppercase; the standard name is \"chrF\" (Popović, 2015). Please use consistent notation.","section":"General"},{"comment":"In the description of prompt-based embeddings, the phrase \"This sentence: \\\"[text]\\\" means in one word: \\\"\" is rendered with escaped quotes in the text; please ensure the prompt is typeset cleanly.","section":"Sections 2 and 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is an empirical study with substantial breadth and a useful comparative resource. The main risk is that the central English-pivot claim is built on PMI cross-target comparability, which is currently validated only circularly and shows a clear failure for one model. The authors should either strengthen the PMI validation or soften the pivot conclusion. The Dali/Belebele circularity and the in-sample model selection are also important but are more local issues that can be addressed with additional analyses or revised claims. The paper has enough merit to warrant a major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful empirical paper and the first, as far as I can tell, to ask whether CLA scores predict MT quality rather than just classification or reading comprehension. It ships a broad comparison of 27 alignment variants, a new few-shot embedding extraction, and a PMI-based translation metric, plus out-of-domain validation on BOUQuET. If the main result holds, English-based alignment becomes a cheap predictor of translation quality, and the practical recommendations (p-wm + ANC) are concrete and testable.\n\nThe strengths are real. The experiment covers three model families, base and instruction-tuned variants, 44 languages, and two translation datasets. Many correlations are strong (0.7-0.9), and the BOUQuET validation suggests the predictive relationship is not just a Flores artifact. The paper is also honest: the Limitations section explicitly says PMI independence cannot be formally established, notes the shared-domain risk, and flags the small size of the related-languages experiment. That candor counts.\n\nThe soft spots, in order of size. First, PMI target-language independence is load-bearing for the en-tgt half of the analysis, and Appendix C validates it circularly: the strong mixed-target correlation with the combined alignment is taken as evidence of independence, but that is the same correlation used to infer the pivot. If PMI retains any target-language-specific offset, en-tgt CLA can correlate with PMI at the language-pair level even if English alignment carries no translation-quality signal. This is the largest worry, and it is not manufactured -- the paper itself concedes full independence cannot be established.\n\nSecond, several choices are tuned in-sample. The alpha/beta weights and the best-of-27 variant selection are fit on Flores before reporting correlations, and the improvement over src-en alone is often tiny. There are no confidence intervals or significance tests. A proper held-out protocol and bootstrap intervals would make the headline numbers trustworthy; right now they are optimistic.\n\nThird, there are small circularities: Dali scores come from Belebele and are then correlated with Belebele accuracy. That inflates that particular cell, though it is not central to the MT claim. Fourth, gemma-3-instruct on Flores is an admitted outlier (PMI-chrF correlation 26.2%) and still enters the main PMI analysis; the paper notes this causes wrong parameter estimation, but a robustness check excluding that cell would help. Fifth, the related-languages experiment has 14 data points and is labeled a side experiment -- treat it as suggestive only.\n\nThe central argument is not a fabrication. The correlations are real, the validation survives out-of-domain, and the limitations are disclosed. But the English-pivot conclusion is weaker than the abstract suggests: the near-zero src-tgt weight could reflect collinearity among alignment scores rather than an internal English step. The paper deserves a serious referee; the path to acceptance is tightening the PMI validation and adding a transparent selection protocol.","headline":"First solid test of whether CLA scores predict MT quality; the English-pivot evidence is real but rests on a partly circular validation of the new PMI metric.","tokens_in":33838,"tokens_out":2029,"would_cite":true,"duration_ms":23867,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"English alignment predicts translation quality as well as direct source-target alignment.","keywords":["cross-lingual alignment","machine translation quality prediction","English pivot language","pointwise mutual information translation metric","sentence embeddings from LLMs","multilingual evaluation","convex combination weights","few-shot embedding extraction"],"falsifier":"Run the same convex-combination analysis with translation quality measured by human ratings or by a reference-based metric corrected for script, instead of PMI. If across 44 target languages the optimal source-target weight rises clearly above zero, or the English-side alignments stop predicting quality, the English-pivot claim would fail. A simpler check: on gemma-3 instruction-tuned, where PMI and chrF disagree, decide which metric reflects true quality; if chrF is right, the near-zero source-target result for that model is an artifact.","tokens_in":1706,"feed_emoji":"🌐","tokens_out":3700,"duration_ms":112635,"temperature":0.7,"pith_summary":"This paper asks whether a model's cross-lingual alignment (CLA) with English is enough to predict how well it will translate between two non-English languages. It compares 27 ways of measuring alignment and introduces a pointwise mutual information (PMI) translation metric that is less tied to the target-language writing system than chrF. Using this metric across 44 languages, the authors find that source-English and English-target alignment predict translation quality comparably to or better than direct source-target alignment, and the direct term gets near-zero weight in an optimal combination. This makes English alignment a cheap prediction tool and supports the idea that LLMs translate through an English internal representation. The exception is closely related language pairs, where direct alignment clearly carries the signal.","feed_headline":"English alignment matches direct as translation quality predictor","feed_subtitle":"Source-to-English and English-to-target carry the signal; direct adds near zero.","key_machinery":"Two devices carry the argument. First, the PMI translation metric: instead of comparing a translation against a reference, it computes $$\\log \\frac{p(y|x)}{p(y)}$$ under the model itself, normalizing the conditional likelihood of a candidate translation by its unconditional likelihood, making scores roughly comparable across target languages. Second, the convex-combination weight analysis: with all three alignment directions available, the paper searches over weights $\\alpha$ and $\\beta$ to maximize correlation between the combined alignment score and translation quality; the optimal weights reveal which alignment direction actually carries the signal.","core_discovery":"The central claim is that the information an LLM's hidden representations carry about translation quality is mostly about how each language aligns with English, not how close source and target are to each other. When translation quality is measured with PMI and predicted by an optimally weighted combination of source-English, English-target, and source-target alignment scores, the source-target weight is consistently near zero or zero for the best-correlating settings across three model families. This pattern survives out-of-domain validation on BOUQuET. For seven pairs of mutually intelligible languages the reverse holds, with source-target alignment dominant. The paper also reproduces earl","pith_inferences":["If the English-pivot interpretation is right, models with more English training data or later instruction tuning might show an even stronger English-target weight; this is directly measurable across more models.","The near-zero source-target weight could partly reflect high correlation among the three alignment directions; a decisive test would use language pairs where source-English and source-target alignment diverge sharply.","PMI, although proposed as a within-model research tool, is a candidate reference-free quality estimate; comparing it to human judgments would test its apparent target-language independence.","The related-language result suggests a boundary for the pivot hypothesis, perhaps a threshold in lexical or typological distance beyond which direct alignment outweighs English alignment."],"forward_implications":["English-side alignment scores suffice as a generation-free predictor of LLM translation quality for model selection and language-pair prioritization.","The English-pivot account is strengthened: translation performance itself is governed by English alignment rather than the source-target relation.","The closely-related-language exception shows the pivot is not mandatory; when translatability is high, direct alignment becomes the better predictor.","Practitioners should prefer position-weighted mean embeddings plus average neuron-wise correlation (ANC), with tokenizer-level alignability as a fallback when hidden states are unavailable.","The PMI metric enables correlation studies across mixed target languages that chrF cannot support, because chrF clusters target languages by character information density."],"supporting_citations":[{"why":"Supplies the English-pivot hypothesis this paper tests: hidden states decode to English-like tokens when translating.","marker":"Wendler et al., 2024"},{"why":"Established that retrieval-based cross-lingual alignment between a language and English predicts LLM performance; the classification baseline this work extends to translation.","marker":"Kargaran et al., 2025"},{"why":"Introduced the Dali retrieval-based alignment score and showed English alignment correlates with multilingual task performance.","marker":"Ravisankar et al., 2026"},{"why":"Provides the xSIM margin-based retrieval procedure underlying the abs, dist, and ratio alignment score variants.","marker":"Artetxe and Schwenk, 2019"},{"why":"Introduced average neuron-wise correlation (ANC), the alignment metric that wins most settings here.","marker":"Del and Fishel, 2022"},{"why":"Proposed the position-weighted mean sentence representation that is the recommended and frequent best representation.","marker":"Muennighoff, 2022"},{"why":"Supplies tokenizer-level alignability via Eflomal, the only non-embedding CLA variant compared.","marker":"Hämmerl et al., 2025"},{"why":"Defines chrF, the reference-based translation quality metric against which the proposed PMI metric is validated.","marker":"Popović, 2015"},{"why":"Supplies the Flores-200 parallel corpus used both to compute alignment scores and to evaluate translations.","marker":"Goyal et al., 2022"},{"why":"Supplies BOUQuET as the out-of-domain validation corpus for translation quality correlations.","marker":"Andrews et al., 2025"}],"fun_headline_variants":["English alignment predicts translation as well as direct measure","Translation quality hinges on English alignment, not source-target","LLMs translate via English pivot, alignment study shows","For LLMs, English is the translation pivot language","Cross-lingual alignment to English predicts translation quality"],"cache_read_input_tokens":35584,"weakest_assumption_plain":"The central claim depends on the PMI metric being a faithful, near-target-language-independent measure of translation quality; the paper itself notes that full independence cannot be formally established and that one model, gemma-3 instruct on Flores, shows a notably weak PMI-chrF correlation of 26.2%.","fun_headline_variants_meta":{"raw":{"variants":["English alignment predicts translation as well as direct measure","Translation quality hinges on English alignment, not source-target","LLMs translate via English pivot, alignment study shows","For LLMs, English is the translation pivot language","Cross-lingual alignment to English predicts translation quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000865,"raw_usage":{"total_tokens":3556,"prompt_tokens":683,"completion_tokens":2873,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":427,"completion_tokens_details":{"reasoning_tokens":2798}},"tokens_in":427,"tokens_out":2873,"duration_ms":23821,"temperature":1.0,"reasoning_tokens":2798,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:58:53.280903+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same convex-combination analysis with translation quality measured by human ratings or by a reference-based metric corrected for script, instead of PMI. If across 44 target languages the optimal source-target weight rises clearly above zero, or the English-side alignments stop predicting quality, the English-pivot claim would fail. A simpler check: on gemma-3 instruction-tuned, where PMI and chrF disagree, decide which metric reflects true quality; if chrF is right, the near-zero source-target result for that model is an artifact.","supporting_citations":[{"cited_title":"MEXA : Multilingual Evaluation of E nglish-Centric LLM s via Cross-Lingual Alignment","cited_arxiv_id":null,"evidence_quote":"Established that retrieval-based cross-lingual alignment between a language and English predicts LLM performance; the classification baseline this work extends to translation."},{"cited_title":"Cross-lingual Similarity of Multilingual Representations Revisited","cited_arxiv_id":null,"evidence_quote":"Introduced average neuron-wise correlation (ANC), the alignment metric that wins most settings here."}],"review_version":1}