REVIEW 54 references
Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
T0 review · reviewed 2026-07-30 · grok-4.5
Pith's one-line read Dynamic MAFC benchmarks still contain 17–29% potentially contaminated post-cut-off claims, and that contamination can inflate Macro-F1 by up to 11 points and reorder models.
desk verdict Solid empirical check that post-cut-off MAFC claims are not automatically clean: ~17–29% still look contaminated, and that can move Macro-F1 by up to ~11 points and scramble rankings. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Evidence-sufficiency contamination score: an LLM is prompted to write a fact-checking article from parametric knowledge alone; evidence items are extracted from both that article and the human oracle article; Hungarian matching with METEOR and two embedding cosine similarities yields a claim-level score; scores above fixed thresholds flag the claim as potentially contaminated.
What would settle it
Re-run the same six models on a fresh post-cut-off claim set after independently human-labeling each claim for whether it is answerable solely from pre-cut-off public knowledge; if the automatically flagged contaminated slice no longer shows higher accuracy/Macro-F1 or different rankings, the central claim fails.
Extended reading notes
Core claim
Dynamic evaluation reduces but does not eliminate contamination risk in multimodal automated fact-checking: 17.09%–29.30% of claims published after model cut-offs remain potentially contaminated under the intersection of three evidence-similarity metrics, and that residual contamination produces statistically significant Macro-F1 inflation (up to 11.34 points) and can reverse system rankings.
Load-bearing premise
High similarity between evidence the model can generate from memory and the evidence human fact-checkers actually used is treated as a valid sign that the model can skip external retrieval.
Editorial extensions
If this is right
- Timestamp filtering alone is insufficient for trustworthy MAFC leaderboards; explicit contamination filters are required.
- Reported gains of agentic MAFC systems on static or lightly filtered dynamic sets may partly reflect memorized evidence rather than retrieval skill.
- Model rankings can reverse once contaminated items are removed, so published orderings on contaminated data are unreliable.
- Future dynamic benchmarks should control label balance and novelty beyond publication date, e.g., by synthesis checks against pre-cut-off knowledge.
- The same detection pipeline can be used as a preprocessing filter to decontaminate existing MAFC datasets.
Reading between the lines
- Any retrieval-augmented or tool-using LLM benchmark that equates ‘after cut-off’ with ‘unseen’ faces the same residual-contamination problem.
- Claims that recombine well-known facts (age rules, old statutes, past news events) will keep leaking into quarterly refreshes unless novelty is checked at the evidence-composition level.
- Conservative lexical-only contamination detectors will systematically under-count risk relative to semantic matching.
- If contamination lets models skip exploration steps, trajectory-level metrics (not just final accuracy) become necessary diagnostics for genuine retrieval ability.
Editorial analysis
A structured set of objections, weighed in public.
Circularity Check
No significant circularity: empirical measurements of contamination and performance are independently defined and externally validated.
full rationale
This is an empirical MAFC evaluation paper, not a first-principles derivation. Contamination is operationalized via Hungarian-matched similarity between evidence extracted from LLM-only fact-checking articles and human oracle articles (Eqs. 1–2, §3.1–3.2), with thresholds fixed by an independent 180-claim human annotation study (τ=25% METEOR; τ=0.6 for embeddings; Appendix D). MAFC Accuracy/Macro-F1 are then measured separately under the DEFAME agent framework on contaminated vs. uncontaminated subsets drawn from the same ClaimReview2025Q4 time window (§4.3–4.4). The performance inflation claim is therefore not forced by the contamination definition: high evidence-overlap labels are not fitted to, nor defined from, the DEFAME scores. Thresholds are not re-tuned on the performance metrics; case studies (Table 4) and trajectory reformulation counts (Table 6) are post-hoc explanations, not algebraic identities. Citations (AVeriTeC pipeline, DEFAME, VERITAS protocol) supply methods and baselines rather than load-bearing uniqueness theorems by the same authors. No self-definitional loop, fitted-input-as-prediction, or renaming of a known identity appears in the derivation chain. Score 0 is appropriate.
Assumptions & free parameters
free parameters (4)
- METEOR contamination threshold τ =
0.25
- Semantic embedding contamination threshold τ =
0.6
- LLM generation temperature =
0.01
- Serper top-k retrieved pages =
3
assumptions (6)
- ad hoc to paper A claim is contaminated for an LLM iff the model’s parametric knowledge can produce evidence sufficiently similar to human oracle evidence (Hungarian-normalized match ≥ τ).
- domain assumption Evidence items extracted by an LLM (default GPT-4o-Mini) from articles preserve the factual content needed for fair oracle vs generated comparison.
- domain assumption Claims published after a stated knowledge cut-off are the right universe for testing residual contamination of “dynamic” evaluation.
- domain assumption Hungarian one-to-one matching with METEOR or embedding cosine is an adequate claim-level similarity aggregator (Eq. 1–2).
- domain assumption IFCN-signatory ClaimReview articles supply reliable oracle verdicts and evidence for English check-worthy claims.
- standard math Bootstrap tests at p<0.05 on Accuracy/Macro-F1 differences indicate statistically meaningful inflation attributable to the contamination split.
invented entities (2)
-
ClaimReview2025Q4 benchmark
-
Claim-level contamination score s(Ẽ,E) and benchmark average S
Cite this review
Pith. "Pith review of Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking." pith.science (2026). https://pith.science/paper/OT4FMRG5
@misc{pith2026260723514,
author = {Pith},
title = {Pith review of: Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking},
year = {2026},
howpublished = {\url{https://pith.science/paper/OT4FMRG5}},
note = {Machine review of arXiv:2607.23514}
}
read the original abstract
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.
Figures
Reference graph
Works this paper leans on
-
[1]
Mubashara Akhtar, Michael Schlichtkrull, Zhijiang Guo, Oana Cocarascu, Elena Simperl, and Andreas Vlachos. 2023. Multimodal Automated Fact-Checking: A Survey. InFindings of EMNLP
2023
-
[2]
Paolo Boldi, Francesco Bonchi, Carlos Castillo, and Sebastiano Vigna. 2011. Query Reformulation Mining: Models, Patterns, and Applications.Information Retrieval 14, 3 (2011), 257–289
2011
-
[3]
Tobias Braun, Mark Rothermel, Marcus Rohrbach, and Anna Rohrbach. 2025. DEFAME: Dynamic Evidence-Based Fact-Checking with Multimodal Experts. In Proc. of ICML
2025
-
[4]
Yuyan Bu, Qiang Sheng, Juan Cao, Peng Qi, Danding Wang, and Jintao Li. 2024. FakingRecipe: Detecting Fake News on Short Video Platforms from the Perspec- tive of Creative Process. InProc. of MM
2024
-
[5]
Grégoire Burel, Martino Mensio, Youri Peskine, Raphael Troncy, Paolo Papotti, and Harith Alani. 2024. CimpleKG: A Continuously Updated Knowledge Graph on Misinformation, Factors and Fact-Checks. InProc. of ISWC
2024
-
[6]
Rui Cao, Zifeng Ding, Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos
-
[7]
Nicoló Fontana, Francesco Corso, Enrico Zuccolotto, and Francesco Pierri. 2025. Evaluating Open-Source Large Language Models for Automated Fact-Checking. arXiv preprint arXiv:2503.05565(2025)
arXiv 2025
-
[8]
Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A Survey on Automated Fact-Checking.Transactions of the Association for Computational Linguistics10 (2022), 178–206
2022
Show all 54 references
-
[9]
Michael Hameleers, Thomas E Powell, Toni GLA Van Der Meer, and Lieke Bos
-
[10]
Haorui He, Yupeng Li, Dacheng Wen, Yang Chen, Reynold Cheng, Donglong Chen, and Francis Lau. 2026. Debating Truth: Debate-Driven Claim Verification with Multiple Large Language Model Agents. InProc. of WWW
2026
-
[11]
Haorui He, Yupeng Li, Bin Benjamin Zhu, Dacheng Wen, Reynold Cheng, and Francis Lau. 2026. Fact2Fiction: Targeted Poisoning Attack to Agentic Fact- Checking System. InProc. of AAAI
2026
-
[12]
Jeff Huang and Efthimis N Efthimiadis. 2009. Analyzing and Evaluating Query Reformulation Strategies in Web Search Logs. InProc. of CIKM
2009
-
[13]
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2025. Live- CodeBench: Holistic and Contamination-Free Evaluation of Large Language Models for Code. InProc. of ICLR
2025
-
[14]
Mohammed Abdul Khaliq, Paul Yu-Chun Chang, Mingyang Ma, Bernhard Pflugfelder, and Filip Miletić. 2024. RAGAR, Your Falsehood Radar: RAG- Augmented Reasoning for Political Fact-Checking Using Multimodal Large Lan- guage Models. InProc. of FEVER
2024
-
[15]
Lev Konstantinovskiy, Oliver Price, Mevan Babakar, and Arkaitz Zubiaga. 2021. Toward Automated Fact-Checking: Developing an Annotation Schema and Bench- mark for Consistent Automated Claim Detection.Digital Threats: Research and Practice2, 2 (2021), 1–16
2021
-
[16]
Yupeng Li, Haorui He, Jin Bai, and Dacheng Wen. 2024. MCFEND: A Multi-Source Benchmark Dataset for Chinese Fake News Detection. InProc. of WWW
2024
-
[17]
Yifeng Luo, Yupeng Li, Dacheng Wen, and Liang Lan. 2024. Message Injection Attack on Rumor Detection under the Black-Box Evasion Setting Using Large Language Model. InProc. of WWW
2024
-
[18]
Martino Mensio and Harith Alani. 2019. MisinfoMe: Who’s Interacting with Misinformation?. InProc. of ISWC
2019
-
[19]
Qiong Nan, Juan Cao, Yongchun Zhu, Yanyan Wang, and Jintao Li. 2021. MD- FEND: Multi-Domain Fake News Detection. InProc. of CIKM
2021
-
[20]
Eryn J Newman, Maryanne Garry, Daniel M Bernstein, Justin Kantner, and D Stephen Lindsay. 2012. Nonprobative Photographs (or Words) Inflate Truthi- ness.Psychonomic Bulletin & Review19, 5 (2012), 969–974
2012
-
[21]
Jingjie Ning, João Coelho, Yibo Kong, Yunfan Long, Bruno Martins, João Ma- galhães, Jamie Callan, and Chenyan Xiong. 2026. Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests. InProc. of SIGIR
2026
-
[22]
Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, and Panagiotis C Petrantonakis. 2024. VERITE: A Robust Benchmark for Multimodal Misinformation Detection Accounting for Unimodal Bias.International Journal of Multimedia Information Retrieval13, 1 (2024), 4
2024
-
[23]
Peng Qi, Yuyan Bu, Juan Cao, Wei Ji, Ruihao Shui, Junbin Xiao, Danding Wang, and Tat-Seng Chua. 2023. FakeSV: A Multimodal Benchmark with Rich Social Context for Fake News Detection on Short Video Platforms. InProc. of AAAI
2023
-
[24]
Dorian Quelle and Alexandre Bovet. 2024. The Perils and Promises of Fact- Checking with Large Language Models.Frontiers in Artificial Intelligence7 (2024)
2024
-
[25]
Mark Rothermel, Marcus Kornmann, Marcus Rohrbach, and Anna Rohrbach
-
[26]
Michael Schlichtkrull, Yulong Chen, Chenxi Whitehouse, Zhenyun Deng, Mubashara Akhtar, Rami Aly, Zhijiang Guo, Christos Christodoulopoulos, Oana Cocarascu, Arpit Mittal, James Thorne, and Andreas Vlachos. 2024. The Auto- mated Verification of Textual Claims (AVeriTeC) Shared T...
2024
-
[27]
Michael Schlichtkrull, Zhijiang Guo, and Andreas Vlachos. 2024. AVeriTeC: A Dataset for Real-World Claim Verification with Evidence from the Web. InProc. of NeurIPS
2024
-
[28]
Qiang Sheng, Juan Cao, Xueyao Zhang, Xirong Li, and Lei Zhong. 2021. Ar- ticle Reranking by Memory-Enhanced Key Sentence Matching for Detecting Previously Fact-Checked Claims. InProc. of ACL
2021
-
[29]
HLE Team. 2025. Humanity’s Last Exam.arXiv preprint arXiv:2501.14249(2025)
2025 arXiv
-
[30]
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal
-
[31]
Andreas Vlachos and Sebastian Riedel. 2014. Fact-Checking: Task Definition and Dataset Construction. InProc. of ACL
2014
-
[32]
Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese
-
[33]
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Sid- dhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, et al
-
[34]
Yuzhuo Xiao, Zeyu Han, Yuhan Wang, and Huaizu Jiang. 2025. XFACTA: Con- temporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs.arXiv preprint arXiv:2508.09999(2025)
2025 arXiv
-
[35]
Barry Menglong Yao, Aditya Shah, Lichao Sun, Jin-Hee Cho, and Lifu Huang
-
[36]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. ReAct: Synergizing Reasoning and Acting in Language Models. InProc. of ICLR
2022
-
[37]
arXiv preprint arXiv:2504.12516(2025)
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents. arXiv preprint arXiv:2504.12516(2025)
2025 arXiv
-
[38]
Yejun Yoon, Jaeyoon Jung, Seunghyun Yoon, and Kunwoo Park. 2025. Hypotheti- cal Documents or Knowledge Leakage? Rethinking LLM-Based Query Expansion. InFindings of ACL
2025
-
[39]
LiveBench: A Challenging, Contamination-Limited LLM Benchmark. In Proc. of ICLR
-
[40]
Linghao Zhang, Shilin He, Chaoyun Zhang, Yu Kang, Bowen Li, Chengxing Xie, Junhao Wang, Maoquan Wang, Yufan Huang, Shengyu Fu, et al. 2025. SWE-Bench Goes Live!. InProc. of NeurIPS
2025
-
[41]
Xueyao Zhang, Juan Cao, Xirong Li, Qiang Sheng, Lei Zhong, and Kai Shu. 2021. Mining Dual Emotion for Fake News Detection. InProc. of WWW
2021
-
[42]
Liwen Zheng, Chaozhuo Li, Xi Zhang, Yu-Ming Shang, Feiran Huang, and Haoran Jia. 2024. Evidence Retrieval Is Almost All You Need for Fact Verification. In Findings of ACL. MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil. He et al., Haorui He, Xinwen Chen, Dacheng Wen, Rey...
2024
-
[44]
Yejun Yoon, Jaeyoon Jung, Seunghyun Yoon, and Kunwoo Park. 2024. HerO at AVeriTeC: The Herd of Open Large Language Models for Verifying Real-World Claims. InProc. of FEVER
2024
-
[46]
Fanrui Zhang, Dian Li, Qiang Zhang, Junxiong Lin, Jiahong Yan, Jiawei Liu, Zheng-Jun Zha, et al. 2025. Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning. InProc. of NeurIPS
2025
-
[50]
Only extract content that appears in the generated text
-
[51]
Evidence must address the main factual assertion(s) made in the claim
-
[52]
Do not infer, summarize, or add information
-
[53]
Do not extract sentences that merely restate the claim
-
[54]
Reason":
Avoid duplication. Fact-Checking Article[FACT-CHECKING ARTICLE] OutputOutput strictly in JSON format: {"Reason": "concise extraction reasoning", "Evidences": [{"Evidence": "... "}]} C Impact of LLM Choice on Evidence Extraction Table 9: Average contamination scores across diff...
-
[2018]
FEVER: A Large-Scale Dataset for Fact Extraction and VERification. In Proc. of NAACL
-
[2020]
A Picture Paints a Thousand Lies? The Effects and Mechanisms of Mul- timodal Disinformation and Rebuttals Disseminated via Social Media.Political Communication37, 2 (2020), 281–301
2020
-
[2023]
End-to-End Multimodal Fact-Checking and Explanation Generation: A Challenging Dataset and Models. InProc. of SIGIR
-
[2025]
AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the Web. InProc. of NeurIPS
-
[2026]
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact- Checking. InProc. of ACL
Reviewed July 30, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.