Pith. sign in

REVIEW 2 major objections 6 minor 74 references

SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval

T0 review · 2 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A 31-team shared task finds that contrastively fine-tuning multilingual embeddings produces the strongest fact-checked claim retrievers and matches translate-to-English pipelines.

desk verdict A genuinely useful shared-task resource paper; the headline finding (crosslingual gap) rests on label-completeness assumptions the paper doesn't audit, so the benchmark deserves peer review with a requested evaluation audit. read the letter →

arxiv 2505.10740 v1 pith:57QMXGYP submitted 2025-05-15 cs.CL cs.IR

classification cs.CLcs.IR
keywords fact-checkedclaimretrievalmultilingualinformationcrosslingualcontrastivefine-tuningtextembeddingssharedtasksuccess@10disinformationdetection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Fact-checking organizations waste effort when they re-check claims that have already been checked in another language. This paper reports a shared task designed to make that problem measurable: given a social media post, retrieve every previously fact-checked claim that matches it, in the same language or across languages, with 31 teams and 52 systems competing on 206,000 fact-checked claims and 28,000 posts. Its main finding is that the winning strategy is not a larger model or more translation, but fine-tuning existing multilingual embedding models with contrastive learning, using in-batch negatives and hard negatives. It also shows that modern multilingual models working on original-language text now perform about as well as the older strategy of translating everything to English and searching there, which matters for low-resource languages. Finally, the results quantify the remaining difficulty: crosslingual retrieval trails monolingual by more than ten points on success@10, and performance on the two unseen languages, Polish and Turkish, is consistently lower.

What carries the argument

The central mechanism is the benchmark itself: a two-track retrieval setup built on a published multilingual dataset of fact-checked claims and social media posts, scored by success@10, meaning whether at least one relevant claim appears in a system's top ten results. The technique that separated the top systems from the baselines is contrastive fine-tuning of embedding models with in-batch negatives, often augmented with hard negatives and followed by re-ranking or weighted voting; the paper identifies this as the most common and most effective strategy among the submitted systems.

What would settle it

Take a random sample of test posts, have multilingual annotators find all matching fact-checked claims across every language in the pool, add any missing links, and recompute success@10; if scores shift differently across language pairs, or if the crosslingual gap narrows or widens, the original labels rather than the systems were driving part of the result.

Watch

Extended reading notes

Core claim

The paper's central claim is that multilingual and crosslingual fact-checked claim retrieval can be organized as a shared task and that, within that task, a clear recipe emerges: fine-tune a multilingual text-embedding model with a contrastive objective and in-batch negatives, optionally add hard negatives, and this outperforms BM25, baseline embedding models, and translate-then-retrieve pipelines while remaining competitive. The best monolingual system reaches 0.9601 average success@10; the best crosslingual system reaches 0.85875. Multilingual models used directly on original languages now give results comparable to English-translation pipelines, a shift from earlier findings on the same underlying dataset. The paper also claims that performance drops by more than ten points in the crosslingual setting and that unseen languages remain the hardest cases, identifying generalization across languages as the open problem.

Load-bearing premise

The evaluation treats the collected post-to-claim links as the complete set of correct answers, although the paper reports that about 15 percent of those links depend on visual information and that many crosslingual connections are missing from the data; if the missing links are distributed unevenly across languages, the reported crosslingual gap could be an artifact of incomplete labels rather than a measure of system ability.

Editorial extensions

If this is right

  • The test set with two previously unseen languages becomes a reusable benchmark for future fact-checked claim retrieval systems.
  • A fact-checking organization can deploy a single contrastively fine-tuned multilingual embedding model, avoiding translation costs, and still match translate-to-English performance on the tested languages.
  • Unseen languages remain the weakest point, so future systems should be evaluated with held-out languages rather than only averaged scores.
  • Crosslingual retrieval remains more than ten points below monolingual, so improvements in cross-language matching are the highest-impact direction.
  • Because combining original text with English translations helps monolingual retrieval but can hurt crosslingual retrieval, recipe choices need to be validated on each track separately.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's claims: because roughly 15 percent of post-claim links depend on images or video, systems that incorporate OCR or image information should be measured separately on that subset, and I would expect them to show outsized gains there.
  • Beyond the paper's claims: if the missing crosslingual labels the paper describes were filled by manual annotation, the more-than-ten-point crosslingual gap could shrink; a targeted annotation study on English-pivot pairs would test this directly.
  • Beyond the paper's claims: the finding that original-language multilingual models now match translate-to-English pipelines suggests a testable extension, measuring zero-shot performance on a language family with no training examples, with and without contrastive fine-tuning on a related language.
  • Beyond the paper's claims: because most crosslingual test pairs involve English as either the post or claim language, the reported crosslingual scores may overstate true diversity; a balanced evaluation controlling for English involvement would give a cleaner estimate.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper reports on SemEval-2025 Task 7, a shared task on multilingual and crosslingual fact-checked claim retrieval. It introduces the task setup, the underlying MultiClaim-based training/development data and a newly collected test set spanning 10 monolingual and 14 crosslingual languages, describes four baselines, and presents the results of 27 monolingual and 28 crosslingual system submissions. The main findings are that contrastive fine-tuning with in-batch negatives was the most common and successful adaptation strategy, that multilingual embedding models are now broadly comparable to translate-to-English pipelines, and that crosslingual retrieval remains markedly harder than monolingual retrieval (a success@10 gap of more than 10 points).

Significance. The shared task is a useful community resource: the dataset is public on Zenodo, the baselines are clearly defined, and the paper aggregates a substantial number of system descriptions, providing guidance for practitioners. The observation that contrastive fine-tuning with in-batch negatives is a robust recipe, and that multilingual models have closed much of the gap with translation-based systems, are valuable empirical signals. The central reliability claim, however, depends on the completeness of the newly collected test-set labels; the paper provides no label-recall audit, and its own description of the collection methodology indicates that missing links are likely. Until that issue is addressed or the crosslingual conclusions are appropriately qualified, the benchmark's quantitative comparisons should be treated with caution.

major comments (2)
  1. [Section 3.1 and Section 3.2] The test-set gold labels are not validated for completeness, and this is load-bearing for the paper's main comparison. The paper reports that in the MultiClaim training/development data, roughly 15% of SMP-claim connections rely on visual information and that the collection procedure misses many potential connections, especially crosslingual ones; the authors mitigate this in training/development by adding 3,351 transitive-closure pairs. No such audit or transitive-closure count is reported for the newly collected test set, which was built with the same methodology. Because success@10 credits a system only if a gold link appears in the top 10, an incompletely labeled test set penalizes systems that retrieve genuinely relevant but unlabeled claims. If the missing links are concentrated in crosslingual pairs, the >10-point difference between tracks in Section 5.2 and the rankings in Table 9 could be partly artifact rather than true retrieval quality. Please report a label-recall audit for the test set (e.g., manual assessment of a sample of top-10 misses per language pair) or explicitly qualify the crosslingual gap and rankings as upper bounds.
  2. [Section 5.2 and Table 4] The crosslingual success@10 is reported only as an aggregate over an extremely imbalanced set of language pairs. Table 4 shows, for example, that the English-Hindi cell contains 2,337 pairs while many cells contain fewer than 10 pairs, and the overall crosslingual average is therefore dominated by a few high-resource, English-centric combinations. The paper should report per-language-pair (or at least per-SMP-language) success@10 to show that the crosslingual gap and the system rankings in Table 9 are not driven by this imbalance, and should discuss how the imbalance interacts with the missing-label concern in Major Comment 1.
minor comments (6)
  1. [Abstract and Introduction] The text says '52 test submissions by 31 teams,' but Tables 8 and 9 list 27 and 28 ranked team rows, totaling 55 team-track submissions. Please reconcile these counts or clarify why some ranked rows are not counted as test submissions.
  2. [Table 2] The columns for numbers of claims and pairs run together (e.g., the English row reads '85,734 145,2875,446 627 574'), making the table very hard to parse; please use separate, clearly labeled columns.
  3. [Tables 8 and 9] Because success@10 is a point estimate over a finite test sample, adjacent ranks differing by less than 0.01 are likely within sampling noise; consider adding confidence intervals or a caveat.
  4. [Section 4] Please state whether the baselines were used out-of-the-box or tuned on the development set, and if tuned, which hyperparameters were selected.
  5. [Section 3.1] The limitations paragraph notes the English-centricity of the crosslingual pairs but should also mention the likely incompleteness of the test-set labels; this is directly relevant to interpreting the results.
  6. [Table 7] The table header is formatted in a way that the column groups are difficult to identify as printed; consider making the column structure explicit with a legend.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: central results rest on newly collected test data and independent participant systems; self-citation is contextual and transparent.

full rationale

The paper's central findings are empirical and evaluated on a newly collected test set (Section 3.1) using submissions from independent teams (Tables 8 and 9), so claims about strategy effectiveness ('the most common strategy to improve base embedding models is to fine-tune them in a contrastive learning manner with in-batch negatives') and about multilingual models being comparable to translate-to-English pipelines are supported by externally produced participant systems. The reuse of MultiClaim (Pikuliak et al., 2023) is transparent dataset construction, not a fitted parameter renamed as a prediction, and the baselines chosen from that prior work are not used to derive the participant rankings. The paper's own caveats about roughly 15% visual-link dependence and missing crosslingual connections are stated limitations about label completeness, not circular derivations; they raise evaluation-validity risk but do not make any claimed result equivalent to its input. No equation is re-derived from itself, and no prediction is statistically forced by a prior fit. Minor self-citation exists (baseline selection, dataset provenance), but it is contextual rather than load-bearing, so the circularity burden is low.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claims do not depend on fitted parameters introduced by this paper; the baselines are public pre-trained models and the participants' hyperparameters are external to the overview. The evaluation does rely on domain assumptions about annotation completeness, the quality of machine translation and OCR, and the adequacy of success@10 as a metric. No new entities are postulated.

assumptions (3)
  • domain assumption The SMP-claim relevance links in the test set are sufficiently complete and correct for success@10 to be an unbiased measure.
    The paper states that roughly 15 percent of links rely on visual content and that many crosslingual connections are missing (Section 3.1), so this assumption is known to be violated to some degree.
  • domain assumption Google Vision API and Google Translate API produce acceptable text transcriptions and translations for posts and claims.
    Used in the data format pipeline (Section 3.1); errors would propagate to both baselines and participant systems.
  • domain assumption The evaluation metric success@10, counting at least one relevant claim in the top ten, captures retrieval quality.
    No precision or ranking position is measured; the paper does not justify this choice.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval." pith.science (2026). https://pith.science/paper/57QMXGYP

@misc{pith2026250510740,
  author       = {Pith},
  title        = {Pith review of: SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/57QMXGYP}},
  note         = {Machine review of arXiv:2505.10740}
}
read the original abstract

The rapid spread of online disinformation presents a global challenge, and machine learning has been widely explored as a potential solution. However, multilingual settings and low-resource languages are often neglected in this field. To address this gap, we conducted a shared task on multilingual claim retrieval at SemEval 2025, aimed at identifying fact-checked claims that match newly encountered claims expressed in social media posts across different languages. The task includes two subtracks: (1) a monolingual track, where social posts and claims are in the same language, and (2) a crosslingual track, where social posts and claims might be in different languages. A total of 179 participants registered for the task contributing to 52 test submissions. 23 out of 31 teams have submitted their system papers. In this paper, we report the best-performing systems as well as the most common and the most effective approaches across both subtracks. This shared task, along with its dataset and participating systems, provides valuable insights into multilingual claim retrieval and automated fact-checking, supporting future research in this field.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

74 extracted references · 43 canonical work pages

  1. [1]

    Mohammad Mahdi Abootorabi, Alireza Ghahramani Kure, Mohammadali Mohammadkhani, Sina Elahimanesh, and Mohammad ali Ali panah. 2025. MultiMind at SemEval-2025 T ask 7: Crosslingual fact-checked claim retrieval via multi-source alignment. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for...

  2. [2]

    Mubashara Akhtar, Rami Aly, Christos Christodoulopoulos, Oana Cocarascu, Zhijiang Guo, Arpit Mittal, Michael Schlichtkrull, James Thorne, and Andreas Vlachos, editors. 2023. https://aclanthology.org/2023.fever-1.0/ Proceedings of the Sixth Fact Extraction and VERification Workshop (FEVER) . Association for Computational Linguistics, Dubrovnik, Croatia

  3. [3]

    Firoj Alam, Alberto Barr \'o n-Cede \ n o, Gullal S Cheema, Gautam Kishore Shahi, Sherzod Hakimov, Maram Hasanain, Chengkai Li, Rub \'e n M \' guez, Hamdy Mubarak, Wajdi Zaghouani, et al. 2023. Overview of the CLEF -2023 CheckThat! L ab T ask 1 on check-worthiness of multimodal and multigenre content

  4. [4]

    Fatma Arslan, Naeemul Hassan, Chengkai Li, and Mark Tremayne. 2020. https://doi.org/10.1609/icwsm.v14i1.7346 A benchmark dataset of check-worthy factual claims . Proceedings of the International AAAI Conference on Web and Social Media, 14(1):821--829

  5. [5]

    Amirmohammad Azadi, Sina Zamani, Mohammadmostafa Rostamkhani, and Sauleh Eetemadi. 2025. Word2winners at SemEval-2025 T ask 7: Multilingual and crosslingual fact-checked claim retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for Computational Linguistics

  6. [6]

    Alberto Barr \'o n-Cede \ n o, Firoj Alam, Tanmoy Chakraborty, Tamer Elsayed, Preslav Nakov, Piotr Przyby a, Julia Maria Stru , Fatima Haouari, Maram Hasanain, Federico Ruggeri, Xingyi Song, and Reem Suwaileh. 2024. The CLEF-2024 CheckThat! L ab: Check-worthiness, subjectivity, persuasion, roles, authorities, and adversarial robustness. In Advances in Inf...

  7. [7]

    Alberto Becerra-Tome and Agustín Conesa. 2025. UPC-HLE at SemEval-2025 T ask 7: Multilingual fact-checked claim retrieval with text embedding models and cross-encoder re-ranking. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for Computational Linguistics

  8. [8]

    Dominik Benchert, Severin Meßlinger, Sven Goller, Jonas Kaiser, Jan Pfister, and Andreas Hotho. 2025. CAIDAS at SemEval-2025 T ask 7: Enriching sparse datasets with LLM -generated content for improved information retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for Computation...

Show all 74 references
  1. [9]

    Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024. http://arxiv.org/abs/2402.03216 Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation

  2. [10]

    Radu Chivereanu and Dan Tufis. 2025. RACAI at SemEval-2025 T ask 7: Efficient adaptation of large language models for multilingual and crosslingual fact-checked claim retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Au...

  3. [11]

    Christos Christodoulopoulos, James Thorne, Andreas Vlachos, Oana Cocarascu, and Arpit Mittal, editors. 2020. https://aclanthology.org/2020.fever-1.0 Proceedings of the Third Workshop on Fact Extraction and VERification (FEVER) . Association for Computational Linguistics, Online

  4. [12]

    European Commission. 2022. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex

  5. [13]

    Gordon V Cormack, Charles LA Clarke, and Stefan Buettcher. 2009. Reciprocal rank fusion outperforms condorcet and individual rank learning methods. In Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, pages 758--759

  6. [14]

    John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, et al. 2024. Aya expanse: Combining research breakthroughs for a new multilingual frontier. arXiv preprint arXiv:2412.04261

  7. [15]

    Prasanna Jagadeesh Devadiga, Arya Suneesh, Pawan Kumar Rajpoot, Bharatdeep Hazarika, and Aditya U. Baliga. 2025. TIFIN I ndia at SemEval -2025 task 7: Harnessing translation to overcome multilingual IR challenges in fact-checked claim retrieval. In Proceedings of the 19th Inte...

  8. [16]

    Alexandru Enache. 2025. UniBuc-AE at SemEval-2025 T ask 7: Training text embedding models for multilingual and crosslingual fact-checked claim retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for C...

  9. [17]

    Ahmet Bahadir Eyuboglu, Bahadir Altun, Mustafa Bora Arslan, Ekrem Sonmezer, and Mucahid Kutlu. 2023. Fight against misinformation on social media: detecting attention-worthy and harmful tweets and verifiable and check-worthy claims. In International Conference of the Cross-Lan...

  10. [18]

    Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. Transactions of the Association for Computational Linguistics, 10:178--206

  11. [19]

    Santiago Lares Harbin and Juan Manuel Pérez. 2025. S houth NLP at SemEval-2025 T ask 7: Multilingual fact-checking retrieval using contrastive learning. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for Co...

  12. [20]

    Muqaddas Haroon, Shaina Ashraf, Ipek Baris, and Lucie Flek. 2025. CAISA at SemEval -2025 T ask 7: Multilingual and crosslingual fact-checked claim retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association f...

  13. [21]

    Andrea Hrckova, Robert Moro, Ivan Srba, Jakub Simko, and Maria Bielikova. 2024. http://arxiv.org/abs/2211.12143 Autonomation, not automation: Activities and needs of fact-checkers as a basis for designing human-centered ai systems

  14. [22]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. https://openreview.net/forum?id=nZeVKeeFYf9 Lo RA : Low-rank adaptation of large language models . In International Conference on Learning Representations

  15. [23]

    Ma \"e l Jullien, Marco Valentino, Hannah Frost, Paul O ' regan, Donal Landers, and Andr \'e Freitas. 2023. https://doi.org/10.18653/v1/2023.semeval-1.307 S em E val-2023 task 7: Multi-evidence natural language inference for clinical trial data . In Proceedings of the 17th Int...

  16. [24]

    Ashkan Kazemi, Kiran Garimella, Devin Gaffney, and Scott Hale. 2021. https://doi.org/10.18653/v1/2021.acl-long.347 Claim matching beyond E nglish to scale global fact-checking . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the ...

  17. [25]

    Mohammed Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, B Suriyaprasaad, G Varun, Sparsh Jain, Anoop Kunchukuttan, Pratyush Kumar, et al. 2024. Indicllmsuite: A blueprint for creating pre-training and fine-tuning datasets for indian languages. ...

  18. [26]

    Lev Konstantinovskiy, Oliver Price, Mevan Babakar, and Arkaitz Zubiaga. 2021. https://doi.org/10.1145/3412869 Toward automated factchecking: Developing an annotation schema and benchmark for consistent automated claim detection . Digital Threats, 2(2)

  19. [27]

    Neema Kotonya and Francesca Toni. 2020. Explainable automated fact-checking: A survey. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5430--5443

  20. [28]

    S Vinod Kumar, Guravana Jothsna, Mummina Poojitha, Kandula Deepika, Indubu Nutana, MVS Ajit, and Konna Lalitendra Swamy. 2024. Verdict prediction using artificial neural networks. In 2024 IEEE International Conference on Blockchain and Distributed Systems Security (ICBDS), pag...

  21. [29]

    Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, et al. 2022. Matryoshka representation learning. Advances in Neural Information Processing Systems, 35:30233--30249

  22. [30]

    Michelle Seng Ah Lee and Jatinder Singh. 2021. https://doi.org/10.1145/3461702.3462572 Risk identification questionnaire for detecting unintended bias in the machine learning development lifecycle . In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIE...

  23. [31]

    Ladislav Lenc, Daniel Cífka, Jiří Martínek, Jakub Šmíd, and Pavel Král. 2025. UWBa at SemEval-2025 T ask 7: Multilingual and crosslingual fact-checked claim retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Ass...

  24. [32]

    Hao Liao, Jiahao Peng, Zhanyi Huang, Wei Zhang, Guanghua Li, Kai Shu, and Xing Xie. 2023. Muser: A multi-step evidence retrieval enhancement framework for fake news detection. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 4461--4472

  25. [33]

    Youzheng Liu, Jiyan Liu, Xiaoman Xu, Taihang Wang, Yimin Wang, and Ye Jiang. 2025. QUST\_NLP at SemEval-2025 T ask 7: A three-stage retrieval framework for monolingual and crosslingual fact-checked claim retrieval. In Proceedings of the 19th International Workshop on Semantic ...

  26. [34]

    Yi-Ju Lu and Cheng-Te Li. 2020. Gcan: Graph-aware co-attention networks for explainable fake news detection on social media. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 505--514

  27. [35]

    Yuheng Mao, Jin Wang, and Xuejie Zhang. 2025. YNU-HPCC at SemEval-2025 T ask 7: Multilingual and cross-lingual fact-checked claim retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for Computational ...

  28. [36]

    Nicholas Micallef, Vivienne Armacost, Nasir Memon, and Sameer Patil. 2022. True or false: Studying the work practices of professional fact-checkers. Proceedings of the ACM on Human-Computer Interaction, 6(CSCW1):1--44

  29. [37]

    Ndapandula Nakashole and Tom Mitchell. 2014. Language-aware truth assessment of fact candidates. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1009--1019

  30. [38]

    Preslav Nakov, Alberto Barr \'o n-Cede \ n o, Giovanni da San Martino, Firoj Alam, Julia Maria Stru , Thomas Mandl, Rub \'e n M \' guez, Tommaso Caselli, Mucahid Kutlu, Wajdi Zaghouani, et al. 2022. Overview of the CLEF --2022 CheckThat! L ab on fighting the COVID-19 infodemic...

  31. [39]

    Preslav Nakov, David Corney, Maram Hasanain, Firoj Alam, Tamer Elsayed, Alberto Barrón-Cedeño, Paolo Papotti, Shaden Shaar, and Giovanni Da San Martino. 2021. https://doi.org/10.24963/ijcai.2021/619 Automated fact-checking for assisting human fact-checkers . In Proceedings of ...

  32. [40]

    Atanu Nayak, Srijani Debnath, Arpan Majumdar, Pritam Pal, and Dipankar Das. 2025. JU\_NLP at SemEval-2025 T ask 7: Leveraging transformer-based models for multilingual & crosslingual fact-checked claim retrieval. In Proceedings of the 19th International Workshop on Semantic Ev...

  33. [41]

    Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022. https://aclanthology.org/2022.emnlp-main.669 Large dual encoders are generalizable retrievers . In Proceedings of the 2022 Confer...

  34. [42]

    Evgenii Nikolaev, Ivan Bondarenko, Islam Aushev, Vasilii Krikunov, Andrei Glinskii, Vasily Konovalov, and Julia Belikova. 2025. FactDebug at SemEval-2025 T ask 7: Hybrid retrieval pipeline for identifying previously fact-checked claims across multiple languages. In Proceedings...

  35. [43]

    Ronghao Pan, Tomás Bernal-Beltrán, José Antonio García-Díaz, and Rafael Valencia-García. 2025. UMUTeam at SemEval-2025 T ask 7: Multilingual fact-checked claim retrieval with XLM-RoBERTa and self-alignment pretraining strategy. In Proceedings of the 19th International Workshop...

  36. [44]

    Rrubaa Panchendrarajan, Rafael Martins Frade, and Arkaitz Zubiaga. 2025. ClaimCatchers at SemEval-2025 T ask 7: Sentence transformers for claim retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for ...

  37. [45]

    Iva Pezo, Allan Hanbury, and Moritz Staudinger. 2025. ipezoTU at SemEval-2025 T ask 7: Hybrid ensemble retrieval for multilingual fact-checking. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for Computatio...

  38. [46]

    Mat \'u s Pikuliak, Ivan Srba, Robert Moro, Timo Hromadka, Timotej Smole n , Martin Meli s ek, Ivan Vykopal, Jakub Simko, Juraj Podrou z ek, and Maria Bielikova. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.1027 Multilingual previously fact-checked claim retrieval . In Pr...

  39. [47]

    Pranshu Rastogi. 2025. f act check AI at SemEval-2025 T ask 7: Multilingual and crosslingual fact-checked claim retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for Computational Linguistics

  40. [48]

    Nils Reimers and Iryna Gurevych. 2019. https://doi.org/10.18653/v1/D19-1410 Sentence- BERT : Sentence embeddings using S iamese BERT -networks . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference...

  41. [49]

    Suprabhat Rijal and Saurav K. Aryal. 2025. Howard University-AI4PC at SemEval-2025 T ask 7: Crosslingual fact-checked claim retrieval-combining zero-shot claim extraction and knn-based classification for multilingual claim matching. In Proceedings of the 19th International Wor...

  42. [50]

    Alex Robertson and Huizhi Liang. 2025. NCL-AR at SemEval-2025 T ask 7: A sieve filtering approach to refute the misinformation within harmful social media posts. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Associati...

  43. [51]

    Stephen Robertson and Hugo Zaragoza. 2009. https://doi.org/10.1561/1500000019 The probabilistic relevance framework: Bm25 and beyond . Foundations and Trends® in Information Retrieval, 3(4):333--389

  44. [52]

    Daniel Russo, Serra Sinem Tekiroğlu, and Marco Guerini. 2023. https://doi.org/10.1162/tacl_a_00601 Benchmarking the generation of fact checking explanations . Transactions of the Association for Computational Linguistics, 11:1250--1264

  45. [53]

    Michael Schlichtkrull, Yulong Chen, Chenxi Whitehouse, Zhenyun Deng, Mubashara Akhtar, Rami Aly, Zhijiang Guo, Christos Christodoulopoulos, Oana Cocarascu, Arpit Mittal, James Thorne, and Andreas Vlachos. 2024. https://aclanthology.org/2024.fever-1.1/ The automated verificatio...

  46. [54]

    Michael Schlichtkrull, Zhijiang Guo, and Andreas Vlachos. 2023. Averitec: A dataset for real-world claim verification with evidence from the web. Advances in Neural Information Processing Systems, 36:65128--65167

  47. [55]

    Tal Schuster, Roei Schuster, Darsh J Shah, and Regina Barzilay. 2020. The limitations of stylometry for detecting machine-generated fake news. Computational Linguistics, 46(2):499--510

  48. [56]

    Shaden Shaar, Firoj Alam, Giovanni Da San Martino, and Preslav Nakov. 2022. https://doi.org/10.18653/v1/2022.findings-naacl.122 The role of context in detecting previously fact-checked claims . In Findings of the Association for Computational Linguistics: NAACL 2022, pages 161...

  49. [57]

    Shaden Shaar, Nikolay Babulkov, Giovanni Da San Martino, and Preslav Nakov. 2020. https://doi.org/10.18653/v1/2020.acl-main.332 That is a known lie: Detecting previously fact-checked claims . In Proceedings of the 58th Annual Meeting of the Association for Computational Lingui...

  50. [58]

    S Suryavardan, Shreyash Mishra, Megha Chakraborty, Parth Patwa, Anku Rani, Aman Chadha, Aishwarya Naresh Reganti, Amitava Das, Amit P Sheth, Manoj Chinnakotla, et al. 2023. Findings of factify 2: multimodal fake news detection. In DE-FACTIFY@ AAAI

  51. [59]

    Shujauddin Syed and Ted Pedersen. 2025. Duluth at SemEval-2025 T ask 7: TF-IDF with optimized vector dimensions for multilingual fact-checked claim retrieval. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association ...

  52. [60]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 a . https://doi.org/10.18653/v1/N18-1074 FEVER : a large-scale dataset for fact extraction and VER ification . In Proceedings of the 2018 Conference of the North A merican Chapter of the Associa...

  53. [61]

    James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal. 2018 b . https://doi.org/10.18653/v1/W18-5501 The fact extraction and VER ification ( FEVER ) shared task . In Proceedings of the First Workshop on Fact Extraction and VER ification (...

  54. [62]

    Andreas Vlachos and Sebastian Riedel. 2014. Fact checking: Task definition and dataset construction. In Proceedings of the ACL 2014 workshop on language technologies and computational social science, pages 18--22

  55. [63]

    Nguyen Vo and Kyumin Lee. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.621 Where are the facts? searching for fact-checked information to alleviate the spread of fake news . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP),...

  56. [64]

    Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The spread of true and false news online. Science, 359(6380):1146--1151

  57. [65]

    Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533

  58. [67]

    Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024 b . Multilingual e5 text embeddings: A technical report. arXiv preprint arXiv:2402.05672

  59. [68]

    Nancy X. R. Wang, Diwakar Mahajan, Marina Danilevsky, and Sara Rosenthal. 2021. https://doi.org/10.18653/v1/2021.semeval-1.39 S em E val-2021 task 9: Fact verification and evidence finding for tabular data in scientific documents ( SEM - TAB - FACTS ) . In Proceedings of the 1...

  60. [69]

    Yuqi Wang and Kangshi Wang. 2025. DKE-Research at SemEval-2025 T ask 7: A unified multilingual framework for cross-lingual and monolingual retrieval with efficient language-specific adaptation. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2...

  61. [70]

    Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023. http://arxiv.org/abs/2309.07597 C-pack: Packaged resources to advance general chinese embedding

  62. [71]

    An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Me...

  63. [72]

    Xia Zeng, Amani S Abumansour, and Arkaitz Zubiaga. 2021. Automated fact-checking: A survey. Language and Linguistics Compass, 15(10):e12438

  64. [73]

    Dun Zhang, Jiacheng Li, Ziyang Zeng, and Fulong Wang. 2024. Jasper and stella: distillation of sota embedding models. arXiv preprint arXiv:2412.19048

  65. [74]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...

  66. [75]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.