{"id":"e84cee15-473b-45fd-8d42-e3fc4e1f1c42","arxiv_id":"2501.14721","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A position paper arguing that comparable corpora can go beyond bilingual lexicon induction, toward lexical semantics, multimodal understanding, and diverse perspectives, while cautioning against English-centric translation pivots.","lead":"This keynote paper reviews the history of comparable corpora and proposes new research directions for them. It argues current bilingual benchmarks miss word-sense ambiguity and that English-centric pivoting introduces cultural bias.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MUSE sense-coverage critique rests on a handful of hand-picked examples; a systematic sense-coverage check is needed before the benchmark can be judged to undervalue comparable corpora.","rationale":"The reader's weakest_assumption correctly identifies the representativeness of Tables 5-6 and Table 7 as the load-bearing point. My reading agrees: the paper is a position/keynote piece, and its most consequential empirical assertion, that MUSE underrepresents word-sense ambiguity and therefore undervalues comparable corpora for lexical semantics, is currently supported only by selected examples and a raw count of identical source-target strings. The concern is not that the paper is internally inconsistent; it is that the evidence is anecdotal. The proposed concrete test would settle the issue by measuring sense coverage of MUSE on a random sample of ambiguous words from a sense-aligned resource. If the sampling confirms low coverage, the strongest claim is strengthened; if it does not, the call for a new benchmark loses its empirical motivation. The historical survey and future directions, such as avoiding pivoting via English and using pictures as prompts, are reasonable and appropriately framed as opportunities. No machine-checked proofs or reproducible evaluations are offered, which is consistent with a keynote, but it means the central empirical claim should be treated as conditional. Since the reader already assigned CONDITIONAL with high confidence, my stress-test does not move the verdict.","tokens_in":11245,"tokens_out":3599,"duration_ms":50900,"concrete_test":"Use a sense-aligned resource, such as the Gale et al. (1992) Hansards annotations or BabelNet, to construct a random sample of N=500 English source words whose French translations differ by sense (for example, bank->banque/banc, duty->droit/devoir, sentence->peine/phrase). For each source word, check whether all sense-specific French translations appear anywhere in the MUSE en->fr and fr->en gold dictionaries. Report recall of sense coverage with a bootstrap confidence interval, and compare against a control set of unambiguous words. If sense coverage exceeds about 90%, the bank->banc omissions are outliers and the strongest claim would fail; if it is below about 50%, the paper's benchmark critique is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.6 makes the strongest empirical claim: 'MUSE may not be as effective as older WSD methods for addressing classic challenges with translations of ambiguous words' and then proposes 'room to introduce a new benchmark.' The supporting evidence is (i) Tables 5-6, selected entries where bank->banc is absent, and (ii) Table 7's count that 73,471 of 113,286 en->fr pairs have identical source and target strings. Neither establishes that MUSE systematically lacks sense-distinguishing translations. MUSE dictionaries are built from Wikipedia cross-lingual links, which are article-based rather than sense-based; absence of bank->banc may indicate that French Wikipedia treats banc as the translation of bench, not that the benchmark is biased. To make the load-bearing inference, that MUSE underestimates the value of comparable corpora for lexical semantics, one must show that a substantial fraction of ambiguous source words lack their alternative sense translations. The paper does not supply that distribution; it supplies anecdotes. This is a support gap rather than an internal contradiction, but it is the hinge on which the call for a new benchmark turns. The paper itself hedges with 'may not,' yet the strongest claim as read by the reviewer goes further and treats the gap as established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper, based on a BUCC-2025 keynote, reviews the history of comparable corpora (CC) from parallel corpora and word-sense disambiguation (WSD) to modern bilingual lexicon induction (BLI), and then proposes several future research directions. The historical sections cover the canonical Bank of English-style WSD examples from the Hansards, the distinction between monolingual and bilingual senses, and the limitations of translating English resources into other languages. The paper's main original critique is that the widely used MUSE benchmark for BLI may not be as effective as older WSD methods for handling ambiguous words, because sense-distinguishing translations such as English bank to French banc are missing from the MUSE dictionaries. The paper also suggests that comparable corpora can enable richer lexical semantics, help burst filter bubbles in news and academia, and support transfer learning to 'growth' languages without inappropriate pivoting through English, including using pictures rather than English text as prompts. It concludes by listing open challenges for the community.","tokens_in":11648,"tokens_out":3881,"duration_ms":34644,"significance":"If the MUSE critique is correct, it would imply that the standard BLI benchmark undervalues the contribution of comparable corpora to lexical semantics, motivating the paper's call for a new benchmark. The historical review is accurate and well-cited, and the proposed directions (reverse translation, picture-based prompts, recommending systems as comparable corpora, and filter-bubble awareness) are genuinely useful research agendas from a senior researcher with deep experience in the field. The paper also explicitly makes falsifiable empirical claims about MUSE's sense coverage and about random walks on MUSE dictionaries, and these claims are testable. The main weakness is that the most important empirical claim, the MUSE sense-coverage gap, is supported only by hand-picked examples and one aggregate overlap count, not by a systematic evaluation. This weakens the rhetorical force of the paper's central challenge, though it does not invalidate the broader agenda.","major_comments":[{"comment":"The claim that MUSE 'may not be as effective as older methods for WSD' rests on a small set of manually selected entries, e.g., bank translating to banque but not banc, rather than a systematic sense-coverage analysis. Since this claim is the load-bearing motivation for proposing a new benchmark, the paper should either provide a systematic evaluation (e.g., for all ambiguous English words in the Table 2 set, report how many of their sense-distinguishing translations occur in MUSE) or explicitly frame the critique as an open question. Without such evidence, the examples in Tables 5-6 remain anecdotes, and the aggregate count of self-translating pairs in Table 7 does not directly address sense coverage.","section":"Section 2.6, Tables 5-6"},{"comment":"The statement that 'if we take a random walk over MUSE dictionaries and start from good, such walks will often take us to synonyms, but rarely to antonyms' is a specific empirical claim presented without data, methodology, or citation. Since this observation is used to motivate the proposed 'theory of translation and collocation based on linear algebra and graph theory,' the paper should either supply the quantitative evidence (e.g., precision of retrieving synonyms versus antonyms via random walks) or explicitly label the statement as a conjecture to be tested by future work.","section":"Section 3.1.3, random-walk observation"}],"minor_comments":[{"comment":"The entry for Portuguese in the S2 Abstracts column reads '1,937.959'; this should be '1,937,959' (thousands separator).","section":"Table 9"},{"comment":"The term 'random walk on MUSE dictionaries' is not defined; please specify how the walk is constructed (e.g., alternating between languages, graph nodes as words in both languages, edge weights from dictionary probabilities or embeddings).","section":"Section 3.1.3"},{"comment":"The paper uses the term 'growth languages' without definition; a brief explanation (e.g., languages with many speakers but fewer resources) would make the discussion accessible to a broader audience.","section":"Section 3.1"},{"comment":"There is a typo in 'Sryia'; it should be 'Syria'.","section":"Section 3.3.1"},{"comment":"The noisy-channel notation could be made more precise: define E and F as English and French word sequences, respectively, and explicitly state the prior and likelihood in Eq. (2).","section":"Section 2.1, Eq. (1)-(2)"}],"recommendation":"major_revision","confidential_remarks":"This is a keynote-style position paper, so the referee should calibrate expectations accordingly: it is not a typical empirical contribution. The main issue is that the MUSE sense-coverage critique, which is the most concrete original claim, lacks the systematic evidence needed to be persuasive. The author could either add a small computational analysis of sense coverage across the MUSE dictionaries or soften the claim into an explicit research question. The random-walk observation in Section 3.1.3 is similarly unsupported and should be flagged. The historical and survey content is solid and suited to the venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a keynote position paper, not a research-results paper. Ken Church reviews the history of comparable and parallel corpora, then floats a few research directions. The historical account is accurate and well cited, and the paper is honest that it is an agenda piece rather than a contribution with new results.\n\nThe genuinely interesting bit is the observation in Section 3.1.3: random walks over MUSE dictionaries seem to track synonymy better than PMI does. That is testable and could be a real lead. The filter-bubble argument and the suggestion to translate from the growth language into English rather than out of English are also reasonable and worth discussing.\n\nNow the soft spots. The load-bearing claim is the MUSE critique in Section 2.6: that MUSE may not capture word-sense ambiguity as well as older WSD methods, e.g., bank to banc. The evidence is a handful of hand-picked examples (Tables 5–6) and a count of self-translating pairs (Table 7). That does not establish that MUSE systematically underrepresents senses. The stress-test note is right: MUSE dictionaries come from Wikipedia cross-lingual links, which are article-level. Absence of bank→banc could mean French Wikipedia treats banc as the translation of bench, not that the benchmark is biased. To support the call for a new benchmark, you'd need a systematic sense-coverage evaluation over a representative sample of ambiguous words. The paper hedges with \"may not,\" but it still uses the gap as motivation. That is a support gap, not an internal contradiction.\n\nAlso, the random-walk observation in Section 3.1.3 is stated without data or a citation. For a keynote that is forgivable, but it is not a result.\n\nI agree with the reader's conditional verdict. The paper is a reasonable community call-to-action, and the historical survey deserves credit. But the strongest claim is not backed by the evidence shown. The paper would benefit from either a systematic MUSE sense-coverage analysis or a clearer framing of the gap as a hypothesis for future work.\n\nWho is this for? People working on bilingual lexicon induction, benchmark design, or low-resource NLP who want a quick historical grounding and a few provocative ideas. It is a good keynote to put in front of a workshop audience, but as a journal paper it needs more rigor. I'd send it to review as a position paper, with the expectation that the empirical claims be either supported or explicitly softened before publication.","headline":"A readable keynote-style position paper on comparable corpora: the historical review is solid, but the strongest empirical claim about MUSE's sense-coverage gap rests on anecdotes rather than systematic evidence.","tokens_in":11989,"tokens_out":1854,"would_cite":false,"duration_ms":18689,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MUSE benchmark misses word senses older methods caught","keywords":["comparable corpora","bilingual lexicon induction","word sense disambiguation","MUSE benchmark","lexical semantics","transfer learning","filter bubbles","multimodal"],"falsifier":"Compute the sense-coverage rate of MUSE by aligning its dictionary pairs to an independent sense inventory (e.g., BabelNet) and measuring how many pairs omit a translation that appears in a parallel corpus for the same source word; if the omission rate is near zero, or if supplying the missing senses does not improve downstream translation quality, the paper's central critique fails.","tokens_in":11043,"feed_emoji":"🌐","tokens_out":6441,"duration_ms":49989,"temperature":0.7,"pith_summary":"This keynote-style paper argues that comparable corpora—text collections in different languages covering similar topics without being translations—remain an underused resource because benchmarks aim too low. The author claims the widely used MUSE benchmark for bilingual lexicon induction skips the word-sense ambiguity that earlier work on parallel corpora explicitly handled, making distributional methods look stronger than they are. The paper then maps four future directions: deeper lexical semantics, contrastive analysis, multimodal comparison, and using diverse perspectives to break filter bubbles.","feed_headline":"MUSE benchmark misses word senses older methods caught","feed_subtitle":"A position paper argues comparable corpora do more than word matching—if benchmarks test ambiguity.","key_machinery":"The argument operates by contrasting two ways of extracting a bilingual lexicon: using a parallel corpus like the Canadian Hansards, where an ambiguous English word's French translation labels its sense (bank as 'banque' for money, 'banc' for river), and using the MUSE benchmark's gold dictionaries, which list single-word translation pairs. The MUSE tables (5-7) expose gaps and a high rate of self-translation, providing the evidence that the benchmark dodges the very ambiguity older WSD methods confronted. This contrast, not a new algorithm, is the device that carries the paper's claim.","core_discovery":"The paper's central claim is that the MUSE benchmark, the standard test for inferring translation lexicons from monolingual corpora, systematically underrepresents word-sense ambiguity. Tables comparing MUSE entries with older Hansard-based sense inventories show that pairs like English 'bank' to French 'banc' are missing, and that most MUSE dictionary pairs are self-translations, so the benchmark rewards words that look alike across languages rather than words whose translation depends on context. Consequently, the paper argues, MUSE's leaderboard understates the contribution comparable corpora can make to lexical semantics, and the community should build benchmarks that test harder, context-dependent translations.","pith_inferences":["We infer that the self-translation bias in MUSE may extend to related multilingual embedding evaluations, so a systematic sense-coverage audit of such benchmarks would be a natural first test of this paper's argument.","We infer that back-translation through MUSE-like dictionaries, often used to augment low-resource data, may inherit the same sense-blindness and produce translations that are accurate for cognates but wrong in context.","We infer that the filter-bubble critique pushes beyond bias removal: rather than scrubbing corpora to one neutral view, future multilingual systems could be designed to generate multiple culturally situated perspectives on the same event, which is a testable objective."],"forward_implications":["If the MUSE gap is real, then bilingual lexicon induction results on MUSE overstate how well distributional methods separate word senses, and the case for comparable corpora in lexical semantics is stronger than the benchmark suggests.","A new benchmark that includes ambiguous translation pairs like bank/banc would give the community a harder test and better demonstrate what comparable corpora can contribute.","Transfer learning to growth languages should start from target-language documents and translate into English, or use vector-space similarity, to avoid imposing American perspectives.","Comparable corpora can be applied to clustering academic papers and to multimodal captioning with culturally diverse labels, beyond simple word matching."],"supporting_citations":[{"why":"Introduces the MUSE benchmark and unsupervised machine translation; supplies the benchmark under critique.","marker":"Lample et al., 2017"},{"why":"Introduces word translation without parallel data and the MUSE gold dictionaries; the other half of the benchmark being critiqued.","marker":"Conneau et al., 2017"},{"why":"Uses a parallel corpus (Hansards) to produce sense-labeled data for WSD; the older standard the paper contrasts with MUSE.","marker":"Gale et al., 1992"},{"why":"Raises the original challenge of word-sense ambiguity in machine translation; sets the stakes for WSD.","marker":"Bar-Hillel, 1960"},{"why":"Introduces the term and early methods of comparable corpora; defines the resource the paper champions.","marker":"Fung and Church, 1994"},{"why":"Early work on identifying word translations in non-parallel texts; another foundational reference for comparable corpora.","marker":"Rapp, 1995"},{"why":"Connects Word2Vec to PMI factorization; used to discuss limitations of distributional methods for lexical semantics.","marker":"Levy and Goldberg, 2014"}],"fun_headline_variants":["MUSE benchmark rewards lookalike words, misses senses","MUSE benchmark undercuts comparable corpora's real value","Why MUSE misses context-dependent translations","Comparable corpora need harder benchmarks than MUSE","MUSE's blind spot: word senses older methods caught"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that MUSE underrepresents word-sense ambiguity rests on a handful of hand-picked examples and one aggregate count of self-translations; if those examples are not representative of the full 113k-pair benchmark, the argued gap may not hold.","fun_headline_variants_meta":{"raw":{"variants":["MUSE benchmark rewards lookalike words, misses senses","MUSE benchmark undercuts comparable corpora's real value","Why MUSE misses context-dependent translations","Comparable corpora need harder benchmarks than MUSE","MUSE's blind spot: word senses older methods caught"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000619,"raw_usage":{"total_tokens":2761,"prompt_tokens":723,"completion_tokens":2038,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":339,"completion_tokens_details":{"reasoning_tokens":1962}},"tokens_in":339,"tokens_out":2038,"duration_ms":17706,"temperature":1.0,"reasoning_tokens":1962,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:51:18.436654+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the sense-coverage rate of MUSE by aligning its dictionary pairs to an independent sense inventory (e.g., BabelNet) and measuring how many pairs omit a translation that appears in a parallel corpus for the same source word; if the omission rate is near zero, or if supplying the missing senses does not improve downstream translation quality, the paper's central critique fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Uses a parallel corpus (Hansards) to produce sense-labeled data for WSD; the older standard the paper contrasts with MUSE."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Raises the original challenge of word-sense ambiguity in machine translation; sets the stakes for WSD."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the term and early methods of comparable corpora; defines the resource the paper champions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Early work on identifying word translations in non-parallel texts; another foundational reference for comparable corpora."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Connects Word2Vec to PMI factorization; used to discuss limitations of distributional methods for lexical semantics."}],"review_version":1}