The XNote dataset and LVLM benchmarks demonstrate that current models face significant challenges in generating accurate, grounded Community Notes for image-based contextual deception.
Mmfakebench: A mixed-source mul- timodal misinformation detection benchmark for lvlms
8 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 8representative citing papers
Introduces claim-conditioned re-scoring (SIFT) and warranted supports proportion (WSP) metric, reporting accuracy recovery up to 27.6 points and WSP calibration at AUC 0.92 on FEVER, SciFact and other benchmarks.
MOMENTA is a unified multimodal architecture using modality-specific experts, bidirectional co-attention, discrepancy detection, temporal drift-momentum aggregation, and domain-adversarial training that reports strong consistent performance on Fakeddit, MMCoVaR, Weibo, and XFacta.
RW-Post is an auditable benchmark linking social media posts to evidence from human fact-check articles for evaluating multimodal AI fact-checking across different evidence regimes.
RW-Post is an auditable text-image benchmark for real-world multimodal fact-checking that links posts to evidence traces from human fact-check articles and includes the AgentFact baseline for evaluation.
MERIT achieves 81.65% F1 on MMFakeBench for multimodal misinformation detection via a four-module framework, outperforming zero-shot baselines like GPT-4V with MMD-Agent at 74.0% F1, with gains attributed to architectural design.
CRAVE is a new framework that clusters retrieved text and image evidence into narratives and uses an LLM judge to produce explained fact-checking verdicts.
FakeVLM-R1 combines GRPO reinforcement learning with critical-thinking CoT and a physics-annotated FakeClue++ dataset to reach claimed SOTA synthetic image detection while reducing over-rejection of real images.
citing papers explorer
-
XNote: Benchmarking Automated Community Notes Generation for Image-based Contextual Deception
The XNote dataset and LVLM benchmarks demonstrate that current models face significant challenges in generating accurate, grounded Community Notes for image-based contextual deception.
-
The Warrant Gap: Claim-Conditioned Re-scoring for Fact-Checking
Introduces claim-conditioned re-scoring (SIFT) and warranted supports proportion (WSP) metric, reporting accuracy recovery up to 27.6 points and WSP calibration at AUC 0.92 on FEVER, SciFact and other benchmarks.
-
MOMENTA: Mixture-of-Experts Over Multimodal Embeddings with Neural Temporal Aggregation for Misinformation Detection
MOMENTA is a unified multimodal architecture using modality-specific experts, bidirectional co-attention, discrepancy detection, temporal drift-momentum aggregation, and domain-adversarial training that reports strong consistent performance on Fakeddit, MMCoVaR, Weibo, and XFacta.
-
RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild
RW-Post is an auditable benchmark linking social media posts to evidence from human fact-check articles for evaluating multimodal AI fact-checking across different evidence regimes.
-
RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild
RW-Post is an auditable text-image benchmark for real-world multimodal fact-checking that links posts to evidence traces from human fact-check articles and includes the AgentFact baseline for evaluation.
-
MERIT: Modular Framework for Multimodal Misinformation Detection with Web-Grounded Reasoning
MERIT achieves 81.65% F1 on MMFakeBench for multimodal misinformation detection via a four-module framework, outperforming zero-shot baselines like GPT-4V with MMD-Agent at 74.0% F1, with gains attributed to architectural design.
-
Fact-Checking with Contextual Narratives: Leveraging Retrieval-Augmented LLMs for Social Media Analysis
CRAVE is a new framework that clusters retrieved text and image evidence into narratives and uses an LLM judge to produce explained fact-checking verdicts.
-
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
FakeVLM-R1 combines GRPO reinforcement learning with critical-thinking CoT and a physics-annotated FakeClue++ dataset to reach claimed SOTA synthetic image detection while reducing over-rejection of real images.