Pith. sign in

REVIEW 54 references

Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

T0 review · reviewed 2026-07-30 · grok-4.5

Pith's one-line read Dynamic MAFC benchmarks still contain 17–29% potentially contaminated post-cut-off claims, and that contamination can inflate Macro-F1 by up to 11 points and reorder models.

desk verdict Solid empirical check that post-cut-off MAFC claims are not automatically clean: ~17–29% still look contaminated, and that can move Macro-F1 by up to ~11 points and scramble rankings. read the letter →

arxiv 2607.23514 v1 pith:OT4FMRG5 submitted 2026-07-26 cs.CL cs.AIcs.MM

classification cs.CLcs.AIcs.MM
keywords multimodalautomatedfact-checkingdynamicevaluationbenchmarkcontaminationknowledgecut-offevidencesufficiencyLLMClaimReviewAVeriTeC
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Multimodal automated fact-checking systems are supposed to retrieve and reason over fresh external evidence, yet static benchmarks are full of old claims that large language models already know. Dynamic benchmarks try to fix this by using only claims published after a model’s knowledge cut-off, on the assumption that those claims are unseen. This paper tests that assumption on the static AVeriTeC set and a new Q4-2025 ClaimReview set. It finds that dynamic collection lowers contamination but does not remove it: between 17% and 29% of post-cut-off claims still look verifiable from the model’s internal knowledge alone. Many “new” claims are simply restatements or recombinations of pre-cut-off public facts. When systems are scored on the contaminated slice versus the clean slice, Macro-F1 can jump by as much as 11 points and model rankings change. The authors therefore re-score six leading models on a strictly decontaminated subset and show that even the best still sits below 56% Macro-F1.

What carries the argument

Evidence-sufficiency contamination score: an LLM is prompted to write a fact-checking article from parametric knowledge alone; evidence items are extracted from both that article and the human oracle article; Hungarian matching with METEOR and two embedding cosine similarities yields a claim-level score; scores above fixed thresholds flag the claim as potentially contaminated.

What would settle it

Re-run the same six models on a fresh post-cut-off claim set after independently human-labeling each claim for whether it is answerable solely from pre-cut-off public knowledge; if the automatically flagged contaminated slice no longer shows higher accuracy/Macro-F1 or different rankings, the central claim fails.

Watch

Extended reading notes

Core claim

Dynamic evaluation reduces but does not eliminate contamination risk in multimodal automated fact-checking: 17.09%–29.30% of claims published after model cut-offs remain potentially contaminated under the intersection of three evidence-similarity metrics, and that residual contamination produces statistically significant Macro-F1 inflation (up to 11.34 points) and can reverse system rankings.

Load-bearing premise

High similarity between evidence the model can generate from memory and the evidence human fact-checkers actually used is treated as a valid sign that the model can skip external retrieval.

Editorial extensions

If this is right

  • Timestamp filtering alone is insufficient for trustworthy MAFC leaderboards; explicit contamination filters are required.
  • Reported gains of agentic MAFC systems on static or lightly filtered dynamic sets may partly reflect memorized evidence rather than retrieval skill.
  • Model rankings can reverse once contaminated items are removed, so published orderings on contaminated data are unreliable.
  • Future dynamic benchmarks should control label balance and novelty beyond publication date, e.g., by synthesis checks against pre-cut-off knowledge.
  • The same detection pipeline can be used as a preprocessing filter to decontaminate existing MAFC datasets.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Any retrieval-augmented or tool-using LLM benchmark that equates ‘after cut-off’ with ‘unseen’ faces the same residual-contamination problem.
  • Claims that recombine well-known facts (age rules, old statutes, past news events) will keep leaking into quarterly refreshes unless novelty is checked at the evidence-composition level.
  • Conservative lexical-only contamination detectors will systematically under-count risk relative to semantic matching.
  • If contamination lets models skip exploration steps, trajectory-level metrics (not just final accuracy) become necessary diagnostics for genuine retrieval ability.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical measurements of contamination and performance are independently defined and externally validated.

full rationale

This is an empirical MAFC evaluation paper, not a first-principles derivation. Contamination is operationalized via Hungarian-matched similarity between evidence extracted from LLM-only fact-checking articles and human oracle articles (Eqs. 1–2, §3.1–3.2), with thresholds fixed by an independent 180-claim human annotation study (τ=25% METEOR; τ=0.6 for embeddings; Appendix D). MAFC Accuracy/Macro-F1 are then measured separately under the DEFAME agent framework on contaminated vs. uncontaminated subsets drawn from the same ClaimReview2025Q4 time window (§4.3–4.4). The performance inflation claim is therefore not forced by the contamination definition: high evidence-overlap labels are not fitted to, nor defined from, the DEFAME scores. Thresholds are not re-tuned on the performance metrics; case studies (Table 4) and trajectory reformulation counts (Table 6) are post-hoc explanations, not algebraic identities. Citations (AVeriTeC pipeline, DEFAME, VERITAS protocol) supply methods and baselines rather than load-bearing uniqueness theorems by the same authors. No self-definitional loop, fitted-input-as-prediction, or renaming of a known identity appears in the derivation chain. Score 0 is appropriate.

Assumptions & free parameters 4 free parameters · 6 assumptions · 2 invented entities

The load-bearing claim rests on an operational contamination definition imported from AVeriTeC’s evidence-sufficiency pipeline, two hand-chosen similarity families with calibrated thresholds, and the modeling choice that parametric evidence overlap implies bypass of retrieval—the core MAFC skill. No new physical entities; free parameters are the decision thresholds and experimental knobs (temperature, top-k search). Background assumptions include standard IR matching (Hungarian assignment), embedding cosine as semantic equivalence, and IFCN/ClaimReview as ground-truth fact-check sources.

free parameters (4)
  • METEOR contamination threshold τ = 0.25
    Binary contaminated/uncontaminated label uses τ=25% following AVeriTeC default; directly controls reported contamination proportions and subset splits.
  • Semantic embedding contamination threshold τ = 0.6
    Chosen via threshold sweep on 180 human-labeled claims to maximize agreement; set to 0.6 for both Gemma-Emb-0.3B and Qwen3-Emb-0.6B.
  • LLM generation temperature = 0.01
    Fixed at 0.01 for article generation and agent runs to encourage determinism; affects sampled evidence text and thus similarity scores.
  • Serper top-k retrieved pages = 3
    DEFAME evaluation retrieves top three web/image results per query; influences measured MAFC accuracy independent of contamination labels.
assumptions (6)
  • ad hoc to paper A claim is contaminated for an LLM iff the model’s parametric knowledge can produce evidence sufficiently similar to human oracle evidence (Hungarian-normalized match ≥ τ).
    Definition in §3.1–3.2; operationalizes contamination for MAFC as evidence sufficiency rather than mere string membership in pretraining data.
  • domain assumption Evidence items extracted by an LLM (default GPT-4o-Mini) from articles preserve the factual content needed for fair oracle vs generated comparison.
    Step 2 of the pipeline; partially stress-tested in Appendix C with GPT-5-Nano.
  • domain assumption Claims published after a stated knowledge cut-off are the right universe for testing residual contamination of “dynamic” evaluation.
    Motivates ClaimReview2025Q4 construction relative to GPT-5.2’s Aug 2025 cut-off (§3.3).
  • domain assumption Hungarian one-to-one matching with METEOR or embedding cosine is an adequate claim-level similarity aggregator (Eq. 1–2).
    Imported from AVeriTeC protocol; normalization by |E| assumed not to distort contamination ranking.
  • domain assumption IFCN-signatory ClaimReview articles supply reliable oracle verdicts and evidence for English check-worthy claims.
    Data inclusion filter in §3.3 for the dynamic benchmark.
  • standard math Bootstrap tests at p<0.05 on Accuracy/Macro-F1 differences indicate statistically meaningful inflation attributable to the contamination split.
    Used in Table 5 and trajectory Table 6 significance markers.
invented entities (2)
  • ClaimReview2025Q4 benchmark
    purpose: Provide a post-cut-off English IFCN claim set (901 claims after filtering) to simulate dynamic MAFC evaluation.
    Newly curated for this study because VERITAS was not open-source and XFACTA was stale; not claimed as a lasting community benchmark, but is a constructed artifact the results depend on.
  • Claim-level contamination score s(Ẽ,E) and benchmark average S
    purpose: Scalarize evidence overlap between LLM-only and oracle evidence sets for thresholding and reporting.
    Defined in Eqs. (1)–(2) by adapting AVeriTeC matching; the numeric rates 17–29% are properties of this score, not of an external standard.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking." pith.science (2026). https://pith.science/paper/OT4FMRG5

@misc{pith2026260723514,
  author       = {Pith},
  title        = {Pith review of: Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OT4FMRG5}},
  note         = {Machine review of arXiv:2607.23514}
}
read the original abstract

Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.

Figures

Figures reproduced from arXiv: 2607.23514 by the authors.

Figure 1
Figure 1. An overview of contamination in MAFC and our [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. An overview of contamination detection via an evidence sufficiency evaluation pipeline. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Contamination score distributions of Qwen3.5- [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Supporting analyses for threshold selection and [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

54 extracted references · 4 linked inside Pith

  1. [1]

    Mubashara Akhtar, Michael Schlichtkrull, Zhijiang Guo, Oana Cocarascu, Elena Simperl, and Andreas Vlachos. 2023. Multimodal Automated Fact-Checking: A Survey. InFindings of EMNLP

  2. [2]

    Paolo Boldi, Francesco Bonchi, Carlos Castillo, and Sebastiano Vigna. 2011. Query Reformulation Mining: Models, Patterns, and Applications.Information Retrieval 14, 3 (2011), 257–289

  3. [3]

    Tobias Braun, Mark Rothermel, Marcus Rohrbach, and Anna Rohrbach. 2025. DEFAME: Dynamic Evidence-Based Fact-Checking with Multimodal Experts. In Proc. of ICML

  4. [4]

    Yuyan Bu, Qiang Sheng, Juan Cao, Peng Qi, Danding Wang, and Jintao Li. 2024. FakingRecipe: Detecting Fake News on Short Video Platforms from the Perspec- tive of Creative Process. InProc. of MM

  5. [5]

    Grégoire Burel, Martino Mensio, Youri Peskine, Raphael Troncy, Paolo Papotti, and Harith Alani. 2024. CimpleKG: A Continuously Updated Knowledge Graph on Misinformation, Factors and Fact-Checks. InProc. of ISWC

  6. [6]

    Rui Cao, Zifeng Ding, Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos

  7. [7]

    Nicoló Fontana, Francesco Corso, Enrico Zuccolotto, and Francesco Pierri. 2025. Evaluating Open-Source Large Language Models for Automated Fact-Checking. arXiv preprint arXiv:2503.05565(2025)

  8. [8]

    Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A Survey on Automated Fact-Checking.Transactions of the Association for Computational Linguistics10 (2022), 178–206

Show all 54 references
  1. [9]

    Michael Hameleers, Thomas E Powell, Toni GLA Van Der Meer, and Lieke Bos

  2. [10]

    Haorui He, Yupeng Li, Dacheng Wen, Yang Chen, Reynold Cheng, Donglong Chen, and Francis Lau. 2026. Debating Truth: Debate-Driven Claim Verification with Multiple Large Language Model Agents. InProc. of WWW

  3. [11]

    Haorui He, Yupeng Li, Bin Benjamin Zhu, Dacheng Wen, Reynold Cheng, and Francis Lau. 2026. Fact2Fiction: Targeted Poisoning Attack to Agentic Fact- Checking System. InProc. of AAAI

  4. [12]

    Jeff Huang and Efthimis N Efthimiadis. 2009. Analyzing and Evaluating Query Reformulation Strategies in Web Search Logs. InProc. of CIKM

  5. [13]

    Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2025. Live- CodeBench: Holistic and Contamination-Free Evaluation of Large Language Models for Code. InProc. of ICLR

  6. [14]

    Mohammed Abdul Khaliq, Paul Yu-Chun Chang, Mingyang Ma, Bernhard Pflugfelder, and Filip Miletić. 2024. RAGAR, Your Falsehood Radar: RAG- Augmented Reasoning for Political Fact-Checking Using Multimodal Large Lan- guage Models. InProc. of FEVER

  7. [15]

    Lev Konstantinovskiy, Oliver Price, Mevan Babakar, and Arkaitz Zubiaga. 2021. Toward Automated Fact-Checking: Developing an Annotation Schema and Bench- mark for Consistent Automated Claim Detection.Digital Threats: Research and Practice2, 2 (2021), 1–16

  8. [16]

    Yupeng Li, Haorui He, Jin Bai, and Dacheng Wen. 2024. MCFEND: A Multi-Source Benchmark Dataset for Chinese Fake News Detection. InProc. of WWW

  9. [17]

    Yifeng Luo, Yupeng Li, Dacheng Wen, and Liang Lan. 2024. Message Injection Attack on Rumor Detection under the Black-Box Evasion Setting Using Large Language Model. InProc. of WWW

  10. [18]

    Martino Mensio and Harith Alani. 2019. MisinfoMe: Who’s Interacting with Misinformation?. InProc. of ISWC

  11. [19]

    Qiong Nan, Juan Cao, Yongchun Zhu, Yanyan Wang, and Jintao Li. 2021. MD- FEND: Multi-Domain Fake News Detection. InProc. of CIKM

  12. [20]

    Eryn J Newman, Maryanne Garry, Daniel M Bernstein, Justin Kantner, and D Stephen Lindsay. 2012. Nonprobative Photographs (or Words) Inflate Truthi- ness.Psychonomic Bulletin & Review19, 5 (2012), 969–974

  13. [21]

    Jingjie Ning, João Coelho, Yibo Kong, Yunfan Long, Bruno Martins, João Ma- galhães, Jamie Callan, and Chenyan Xiong. 2026. Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests. InProc. of SIGIR

  14. [22]

    Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, and Panagiotis C Petrantonakis. 2024. VERITE: A Robust Benchmark for Multimodal Misinformation Detection Accounting for Unimodal Bias.International Journal of Multimedia Information Retrieval13, 1 (2024), 4

  15. [23]

    Peng Qi, Yuyan Bu, Juan Cao, Wei Ji, Ruihao Shui, Junbin Xiao, Danding Wang, and Tat-Seng Chua. 2023. FakeSV: A Multimodal Benchmark with Rich Social Context for Fake News Detection on Short Video Platforms. InProc. of AAAI

  16. [24]

    Dorian Quelle and Alexandre Bovet. 2024. The Perils and Promises of Fact- Checking with Large Language Models.Frontiers in Artificial Intelligence7 (2024)

  17. [25]

    Mark Rothermel, Marcus Kornmann, Marcus Rohrbach, and Anna Rohrbach

  18. [26]

    Michael Schlichtkrull, Yulong Chen, Chenxi Whitehouse, Zhenyun Deng, Mubashara Akhtar, Rami Aly, Zhijiang Guo, Christos Christodoulopoulos, Oana Cocarascu, Arpit Mittal, James Thorne, and Andreas Vlachos. 2024. The Auto- mated Verification of Textual Claims (AVeriTeC) Shared T...

  19. [27]

    Michael Schlichtkrull, Zhijiang Guo, and Andreas Vlachos. 2024. AVeriTeC: A Dataset for Real-World Claim Verification with Evidence from the Web. InProc. of NeurIPS

  20. [28]

    Qiang Sheng, Juan Cao, Xueyao Zhang, Xirong Li, and Lei Zhong. 2021. Ar- ticle Reranking by Memory-Enhanced Key Sentence Matching for Detecting Previously Fact-Checked Claims. InProc. of ACL

  21. [29]

    HLE Team. 2025. Humanity’s Last Exam.arXiv preprint arXiv:2501.14249(2025)

  22. [30]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal

  23. [31]

    Andreas Vlachos and Sebastian Riedel. 2014. Fact-Checking: Task Definition and Dataset Construction. InProc. of ACL

  24. [32]

    Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese

  25. [33]

    Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Sid- dhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, et al

  26. [34]

    Yuzhuo Xiao, Zeyu Han, Yuhan Wang, and Huaizu Jiang. 2025. XFACTA: Con- temporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs.arXiv preprint arXiv:2508.09999(2025)

  27. [35]

    Barry Menglong Yao, Aditya Shah, Lichao Sun, Jin-Hee Cho, and Lifu Huang

  28. [36]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. ReAct: Synergizing Reasoning and Acting in Language Models. InProc. of ICLR

  29. [37]

    arXiv preprint arXiv:2504.12516(2025)

    BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents. arXiv preprint arXiv:2504.12516(2025)

  30. [38]

    Yejun Yoon, Jaeyoon Jung, Seunghyun Yoon, and Kunwoo Park. 2025. Hypotheti- cal Documents or Knowledge Leakage? Rethinking LLM-Based Query Expansion. InFindings of ACL

  31. [39]

    LiveBench: A Challenging, Contamination-Limited LLM Benchmark. In Proc. of ICLR

  32. [40]

    Linghao Zhang, Shilin He, Chaoyun Zhang, Yu Kang, Bowen Li, Chengxing Xie, Junhao Wang, Maoquan Wang, Yufan Huang, Shengyu Fu, et al. 2025. SWE-Bench Goes Live!. InProc. of NeurIPS

  33. [41]

    Xueyao Zhang, Juan Cao, Xirong Li, Qiang Sheng, Lei Zhong, and Kai Shu. 2021. Mining Dual Emotion for Fake News Detection. InProc. of WWW

  34. [42]

    Liwen Zheng, Chaozhuo Li, Xi Zhang, Yu-Ming Shang, Feiran Huang, and Haoran Jia. 2024. Evidence Retrieval Is Almost All You Need for Fact Verification. In Findings of ACL. MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil. He et al., Haorui He, Xinwen Chen, Dacheng Wen, Rey...

  35. [44]

    Yejun Yoon, Jaeyoon Jung, Seunghyun Yoon, and Kunwoo Park. 2024. HerO at AVeriTeC: The Herd of Open Large Language Models for Verifying Real-World Claims. InProc. of FEVER

  36. [46]

    Fanrui Zhang, Dian Li, Qiang Zhang, Junxiong Lin, Jiahong Yan, Jiawei Liu, Zheng-Jun Zha, et al. 2025. Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning. InProc. of NeurIPS

  37. [50]

    Only extract content that appears in the generated text

  38. [51]

    Evidence must address the main factual assertion(s) made in the claim

  39. [52]

    Do not infer, summarize, or add information

  40. [53]

    Do not extract sentences that merely restate the claim

  41. [54]

    Reason":

    Avoid duplication. Fact-Checking Article[FACT-CHECKING ARTICLE] OutputOutput strictly in JSON format: {"Reason": "concise extraction reasoning", "Evidences": [{"Evidence": "... "}]} C Impact of LLM Choice on Evidence Extraction Table 9: Average contamination scores across diff...

  42. [2018]

    FEVER: A Large-Scale Dataset for Fact Extraction and VERification. In Proc. of NAACL

  43. [2020]

    A Picture Paints a Thousand Lies? The Effects and Mechanisms of Mul- timodal Disinformation and Rebuttals Disseminated via Social Media.Political Communication37, 2 (2020), 281–301

  44. [2023]

    End-to-End Multimodal Fact-Checking and Explanation Generation: A Challenging Dataset and Models. InProc. of SIGIR

  45. [2025]

    AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the Web. InProc. of NeurIPS

  46. [2026]

    VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact- Checking. InProc. of ACL

Pith tools

Reviewed July 30, 2026 · model on record in the stance chip above.