{"id":"8865c7e9-f32d-4811-a15f-5b6ac1d143e4","arxiv_id":"2505.10740","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A shared task evaluation shows that contrastive fine-tuning of multilingual embedding models is the most common and among the most effective approaches for fact-checked claim retrieval.","lead":"This paper reports SemEval-2025 Task 7, a shared competition in which 31 teams built systems to retrieve previously fact-checked claims matching social media posts across ten to fourteen languages. The report describes the new test data, the best-performing systems, and the techniques that worked best, providing a benchmark for multilingual claim retrieval.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated test-set label completeness makes the crosslingual success@10 gap potentially an artifact; a label-recall audit on the test set is needed.","rationale":"The paper's central claim is that Task 7 constitutes a reliable multilingual/crosslingual benchmark and that its results yield insights, including the empirical finding that contrastive fine-tuning with in-batch negatives is the most common strategy and that multilingual models are now comparable to translate-to-English pipelines. For the benchmark to be reliable, success@10 must measure retrieval quality rather than label completeness. The paper itself states that roughly 15% of connections need visual information and that the collection procedure misses many crosslingual links; those statements appear in Section 3.1, but no completeness estimate is given for the newly collected test set. If test gold labels omit relevant claims, systems are penalized for retrieving them, and if the omissions are concentrated in crosslingual or low-resource pairs, the observed crosslingual gap can be inflated. Table 4 shows the test crosslingual pairs are heavily English-centered, so the missing-label risk is plausibly heterogeneous. This concern is more load-bearing than the numeric 52-vs-55 submission discrepancy or the absence of confidence intervals: those are presentational, while label completeness threatens the validity of every headline comparison. The proposed label-recall audit would directly settle whether the concern lands. I therefore keep the reader's CONDITIONAL verdict and do not change the recommendation.","tokens_in":19179,"tokens_out":6960,"duration_ms":72576,"concrete_test":"Run a label-recall audit on a random sample of at least 200 crosslingual test posts, following the manual inspection protocol of Pikuliak et al. (2023): for each post, have annotators independently search the full test claim pool for all relevant claims, including those not in the gold links. Report gold-label recall per language pair, then recompute success@10 for the systems in Table 9 (and the top monolingual systems in Table 8) using the union of gold and newly found relevant claims. If the crosslingual gap and top-5 rankings are unchanged, the incompleteness concern is not material; if the gap narrows substantially or rankings shift, the reported crosslingual comparison is partly an artifact of missing labels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 concedes two facts that directly bear on the paper's headline comparisons: roughly 15% of SMP-claim links depend on visual information, so a text-only retriever cannot succeed on them, and the collection procedure misses \"many potential connections ... especially crosslingual connections.\" The transitive-closure mitigation is described for the MultiClaim training/development data, but the paper reports no label-recall estimate for the newly collected test set. Since success@10 (Section 3.2) gives credit only for a gold link in the top 10, any system that retrieves a genuinely relevant claim absent from the gold labels is scored as a failure. The crosslingual test set is also heavily English-centered (Table 4), so if missing labels are concentrated in crosslingual or low-resource pairs, the reported 10+ point monolingual-vs-crosslingual gap (Section 5.2) and the Table 9 rankings could be partly an artifact of label incompleteness rather than a measure of retrieval capability. The paper does not quantify completeness for the test set, so this assumption is load-bearing for the benchmark-reliability claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports on SemEval-2025 Task 7, a shared task on multilingual and crosslingual fact-checked claim retrieval. It introduces the task setup, the underlying MultiClaim-based training/development data and a newly collected test set spanning 10 monolingual and 14 crosslingual languages, describes four baselines, and presents the results of 27 monolingual and 28 crosslingual system submissions. The main findings are that contrastive fine-tuning with in-batch negatives was the most common and successful adaptation strategy, that multilingual embedding models are now broadly comparable to translate-to-English pipelines, and that crosslingual retrieval remains markedly harder than monolingual retrieval (a success@10 gap of more than 10 points).","tokens_in":19362,"tokens_out":13740,"duration_ms":123357,"significance":"The shared task is a useful community resource: the dataset is public on Zenodo, the baselines are clearly defined, and the paper aggregates a substantial number of system descriptions, providing guidance for practitioners. The observation that contrastive fine-tuning with in-batch negatives is a robust recipe, and that multilingual models have closed much of the gap with translation-based systems, are valuable empirical signals. The central reliability claim, however, depends on the completeness of the newly collected test-set labels; the paper provides no label-recall audit, and its own description of the collection methodology indicates that missing links are likely. Until that issue is addressed or the crosslingual conclusions are appropriately qualified, the benchmark's quantitative comparisons should be treated with caution.","major_comments":[{"comment":"The test-set gold labels are not validated for completeness, and this is load-bearing for the paper's main comparison. The paper reports that in the MultiClaim training/development data, roughly 15% of SMP-claim connections rely on visual information and that the collection procedure misses many potential connections, especially crosslingual ones; the authors mitigate this in training/development by adding 3,351 transitive-closure pairs. No such audit or transitive-closure count is reported for the newly collected test set, which was built with the same methodology. Because success@10 credits a system only if a gold link appears in the top 10, an incompletely labeled test set penalizes systems that retrieve genuinely relevant but unlabeled claims. If the missing links are concentrated in crosslingual pairs, the >10-point difference between tracks in Section 5.2 and the rankings in Table 9 could be partly artifact rather than true retrieval quality. Please report a label-recall audit for the test set (e.g., manual assessment of a sample of top-10 misses per language pair) or explicitly qualify the crosslingual gap and rankings as upper bounds.","section":"Section 3.1 and Section 3.2"},{"comment":"The crosslingual success@10 is reported only as an aggregate over an extremely imbalanced set of language pairs. Table 4 shows, for example, that the English-Hindi cell contains 2,337 pairs while many cells contain fewer than 10 pairs, and the overall crosslingual average is therefore dominated by a few high-resource, English-centric combinations. The paper should report per-language-pair (or at least per-SMP-language) success@10 to show that the crosslingual gap and the system rankings in Table 9 are not driven by this imbalance, and should discuss how the imbalance interacts with the missing-label concern in Major Comment 1.","section":"Section 5.2 and Table 4"}],"minor_comments":[{"comment":"The text says '52 test submissions by 31 teams,' but Tables 8 and 9 list 27 and 28 ranked team rows, totaling 55 team-track submissions. Please reconcile these counts or clarify why some ranked rows are not counted as test submissions.","section":"Abstract and Introduction"},{"comment":"The columns for numbers of claims and pairs run together (e.g., the English row reads '85,734 145,2875,446 627 574'), making the table very hard to parse; please use separate, clearly labeled columns.","section":"Table 2"},{"comment":"Because success@10 is a point estimate over a finite test sample, adjacent ranks differing by less than 0.01 are likely within sampling noise; consider adding confidence intervals or a caveat.","section":"Tables 8 and 9"},{"comment":"Please state whether the baselines were used out-of-the-box or tuned on the development set, and if tuned, which hyperparameters were selected.","section":"Section 4"},{"comment":"The limitations paragraph notes the English-centricity of the crosslingual pairs but should also mention the likely incompleteness of the test-set labels; this is directly relevant to interpreting the results.","section":"Section 3.1"},{"comment":"The table header is formatted in a way that the column groups are difficult to identify as printed; consider making the column structure explicit with a legend.","section":"Table 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serviceable shared-task overview, and the label-completeness issue is the main threat to its central comparison. The participation-count inconsistency in the abstract should be fixed. I have no concerns about citation practice; the authors' reliance on their own MultiClaim paper is transparent and appropriate. The paper fits the scope of the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a competent shared-task overview that earns its place as a resource paper. The genuinely new part is the test set: newly collected posts and claims up to March/May 2024, two unannounced languages (Polish and Turkish), and a fresh leaderboard of 31 teams across monolingual and crosslingual tracks. The dataset is on Zenodo, and 23 system papers back up the leaderboard entries. That is real evidence, and it is more than the prior MultiClaim paper offered. The empirical summary is also useful: contrastive fine-tuning with in-batch negatives (plus hard negatives) was the most common successful strategy, and strong multilingual embedding models now perform roughly on par with translate-to-English pipelines. Those conclusions are consistent with the tables and with the system-paper reports.\n\nThe soft spots are real but not fatal. The biggest one is label completeness in the test set. The paper concedes that roughly 15% of SMP-claim connections depend on visual information and that the collection procedure misses many crosslingual connections. Transitive closure was used to add missing pairs in the training/development data, but there is no report of a label-recall audit for the new test set. Since success@10 only credits a gold label in the top 10, any system retrieving a genuinely relevant but unlabelled claim is scored as a failure. If missing labels concentrate in crosslingual or low-resource pairs—which the collection procedure suggests they do—the paper's headline 10-point gap between monolingual and crosslingual performance could be inflated. The authors do flag the English-centric nature of the crosslingual test, but they do not quantify test-set completeness. A post-hoc label-recall audit, or at least an analysis of the visual-connection share in the test set, would materially strengthen the benchmark claims.\n\nSmaller issues: the abstract says 52 test submissions while Tables 8 and 9 sum to 55 team-track rows. That is likely a reporting difference between unique teams and track entries, but it needs to be stated. Also, all results are point estimates without error bars or significance tests. That is common in shared-task papers, but the paper occasionally makes comparative claims about techniques (e.g., reranking and weighted voting helping more in crosslingual) that are based on small numbers of systems and no uncertainty quantification. The discussion is mostly appropriately hedged, so this is a request for tightening, not a rewrite.\n\nWho is this for? Anyone building multilingual claim-retrieval systems or planning a shared task. The dataset and leaderboard are the contribution, not the analysis. My recommendation: send it to peer review, and ask the authors to add a test-set label-completeness audit and fix the submission-count inconsistency. That would make it a solid reference.","headline":"A genuinely useful shared-task resource paper; the headline finding (crosslingual gap) rests on label-completeness assumptions the paper doesn't audit, so the benchmark deserves peer review with a requested evaluation audit.","tokens_in":19923,"tokens_out":2915,"would_cite":true,"duration_ms":28166,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 31-team shared task finds that contrastively fine-tuning multilingual embeddings produces the strongest fact-checked claim retrievers and matches translate-to-English pipelines.","keywords":["fact-checked claim retrieval","multilingual information retrieval","crosslingual retrieval","contrastive fine-tuning","text embeddings","shared task","success@10","disinformation detection"],"falsifier":"Take a random sample of test posts, have multilingual annotators find all matching fact-checked claims across every language in the pool, add any missing links, and recompute success@10; if scores shift differently across language pairs, or if the crosslingual gap narrows or widens, the original labels rather than the systems were driving part of the result.","tokens_in":18985,"feed_emoji":"🌐","tokens_out":10730,"duration_ms":93376,"temperature":0.7,"pith_summary":"Fact-checking organizations waste effort when they re-check claims that have already been checked in another language. This paper reports a shared task designed to make that problem measurable: given a social media post, retrieve every previously fact-checked claim that matches it, in the same language or across languages, with 31 teams and 52 systems competing on 206,000 fact-checked claims and 28,000 posts. Its main finding is that the winning strategy is not a larger model or more translation, but fine-tuning existing multilingual embedding models with contrastive learning, using in-batch negatives and hard negatives. It also shows that modern multilingual models working on original-language text now perform about as well as the older strategy of translating everything to English and searching there, which matters for low-resource languages. Finally, the results quantify the remaining difficulty: crosslingual retrieval trails monolingual by more than ten points on success@10, and performance on the two unseen languages, Polish and Turkish, is consistently lower.","feed_headline":"Fine-tuned embeddings rival translation for claim retrieval","feed_subtitle":"In a 31-team shared task, contrastive fine-tuning with in-batch negatives made the strongest claim retrievers.","key_machinery":"The central mechanism is the benchmark itself: a two-track retrieval setup built on a published multilingual dataset of fact-checked claims and social media posts, scored by success@10, meaning whether at least one relevant claim appears in a system's top ten results. The technique that separated the top systems from the baselines is contrastive fine-tuning of embedding models with in-batch negatives, often augmented with hard negatives and followed by re-ranking or weighted voting; the paper identifies this as the most common and most effective strategy among the submitted systems.","core_discovery":"The paper's central claim is that multilingual and crosslingual fact-checked claim retrieval can be organized as a shared task and that, within that task, a clear recipe emerges: fine-tune a multilingual text-embedding model with a contrastive objective and in-batch negatives, optionally add hard negatives, and this outperforms BM25, baseline embedding models, and translate-then-retrieve pipelines while remaining competitive. The best monolingual system reaches 0.9601 average success@10; the best crosslingual system reaches 0.85875. Multilingual models used directly on original languages now give results comparable to English-translation pipelines, a shift from earlier findings on the same underlying dataset. The paper also claims that performance drops by more than ten points in the crosslingual setting and that unseen languages remain the hardest cases, identifying generalization across languages as the open problem.","pith_inferences":["Beyond the paper's claims: because roughly 15 percent of post-claim links depend on images or video, systems that incorporate OCR or image information should be measured separately on that subset, and I would expect them to show outsized gains there.","Beyond the paper's claims: if the missing crosslingual labels the paper describes were filled by manual annotation, the more-than-ten-point crosslingual gap could shrink; a targeted annotation study on English-pivot pairs would test this directly.","Beyond the paper's claims: the finding that original-language multilingual models now match translate-to-English pipelines suggests a testable extension, measuring zero-shot performance on a language family with no training examples, with and without contrastive fine-tuning on a related language.","Beyond the paper's claims: because most crosslingual test pairs involve English as either the post or claim language, the reported crosslingual scores may overstate true diversity; a balanced evaluation controlling for English involvement would give a cleaner estimate."],"forward_implications":["The test set with two previously unseen languages becomes a reusable benchmark for future fact-checked claim retrieval systems.","A fact-checking organization can deploy a single contrastively fine-tuned multilingual embedding model, avoiding translation costs, and still match translate-to-English performance on the tested languages.","Unseen languages remain the weakest point, so future systems should be evaluated with held-out languages rather than only averaged scores.","Crosslingual retrieval remains more than ten points below monolingual, so improvements in cross-language matching are the highest-impact direction.","Because combining original text with English translations helps monolingual retrieval but can hurt crosslingual retrieval, recipe choices need to be validated on each track separately."],"supporting_citations":[{"why":"Supplies the multilingual dataset of fact-checked claims and social media posts that the shared task is built on, plus earlier baseline results that motivate the comparison.","marker":"Pikuliak et al. (2023)"},{"why":"Defines claim matching beyond English, establishing the multilingual scope that the task extends to more languages and language pairs.","marker":"Kazemi et al. (2021)"},{"why":"Frames claim retrieval as searching for fact-checked information and provides the earlier English-centric formulation.","marker":"Vo and Lee, 2020"},{"why":"Provides the BM25 retrieval baseline used as an established non-neural comparison.","marker":"Robertson and Zaragoza (2009)"},{"why":"Supplies the GTR-T5-Large English embedding baseline, the strongest retriever in earlier work, which the multilingual systems now match or beat.","marker":"Ni et al. (2022)"},{"why":"Supplies the multilingual E5 baseline model that several top systems fine-tune with contrastive learning and in-batch negatives.","marker":"Wang et al. (2024a)"}],"fun_headline_variants":["Fine-tuned embeddings beat translation for claim lookup","Contrastive tuning tops multilingual claim retrieval","Multilingual embeddings outdo translate-then-retrieve for claims","Crosslingual claim search: fine-tuned embeddings lead","In-batch negatives power best claim retriever"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation treats the collected post-to-claim links as the complete set of correct answers, although the paper reports that about 15 percent of those links depend on visual information and that many crosslingual connections are missing from the data; if the missing links are distributed unevenly across languages, the reported crosslingual gap could be an artifact of incomplete labels rather than a measure of system ability.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuned embeddings beat translation for claim lookup","Contrastive tuning tops multilingual claim retrieval","Multilingual embeddings outdo translate-then-retrieve for claims","Crosslingual claim search: fine-tuned embeddings lead","In-batch negatives power best claim retriever"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000329,"raw_usage":{"total_tokens":1821,"prompt_tokens":918,"completion_tokens":903,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":827}},"tokens_in":534,"tokens_out":903,"duration_ms":8654,"temperature":1.0,"reasoning_tokens":827,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:04:07.958456+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of test posts, have multilingual annotators find all matching fact-checked claims across every language in the pool, add any missing links, and recompute success@10; if scores shift differently across language pairs, or if the crosslingual gap narrows or widens, the original labels rather than the systems were driving part of the result.","supporting_citations":[],"review_version":1}