Pith. sign in

REVIEW 8 cited by

LEMMA: Towards LVLM-Enhanced Multimodal Misinformation Detection with External Knowledge Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.11943 v2 pith:5GPQXTR7 submitted 2024-02-19 cs.CL

LEMMA: Towards LVLM-Enhanced Multimodal Misinformation Detection with External Knowledge Augmentation

classification cs.CL
keywords lvlmmisinformationdetectionknowledgemultimodalreasoningexternallemma
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The rise of multimodal misinformation on social platforms poses significant challenges for individuals and societies. Its increased credibility and broader impact compared to textual misinformation make detection complex, requiring robust reasoning across diverse media types and profound knowledge for accurate verification. The emergence of Large Vision Language Model (LVLM) offers a potential solution to this problem. Leveraging their proficiency in processing visual and textual information, LVLM demonstrates promising capabilities in recognizing complex information and exhibiting strong reasoning skills. In this paper, we first investigate the potential of LVLM on multimodal misinformation detection. We find that even though LVLM has a superior performance compared to LLMs, its profound reasoning may present limited power with a lack of evidence. Based on these observations, we propose LEMMA: LVLM-Enhanced Multimodal Misinformation Detection with External Knowledge Augmentation. LEMMA leverages LVLM intuition and reasoning capabilities while augmenting them with external knowledge to enhance the accuracy of misinformation detection. Our method improves the accuracy over the top baseline LVLM by 7% and 13% on Twitter and Fakeddit datasets respectively.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild

    cs.MM 2026-05 unverdicted novelty 7.0

    RW-Post is a new post-aligned text-image benchmark with auditable annotations from real fact-checks that reveals current LVLMs struggle with faithful evidence grounding but improve under evidence-bounded evaluation.

  2. Verification-Notebook Learning for Source-Aware Multimodal Misinformation Detection

    cs.AI 2026-07 conditional novelty 6.0

    Verification-Notebook Learning distills labeled multimodal verification experience into a compact fixed notebook that lifts frozen-LVLM source-aware misinformation detection above prompting, cases, and agents.

  3. Detecting AI-Generated Video: A Vision-Language Dual-View Survey

    cs.CV 2026-07 conditional novelty 6.0

    AIGC-V detection should be treated as factual fidelity verification and organized by a four-layer vision-language dual-view taxonomy spanning cues, motion, cross-modal consistency, and world-level reasoning.

  4. RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild

    cs.AI 2025-12 unverdicted novelty 6.0

    RW-Post is an auditable benchmark linking social media posts to evidence from human fact-check articles for evaluating multimodal AI fact-checking across different evidence regimes.

  5. RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild

    cs.MM 2026-05 unverdicted novelty 5.0

    RW-Post is an auditable text-image benchmark for real-world multimodal fact-checking that links posts to evidence traces from human fact-check articles and includes the AgentFact baseline for evaluation.

  6. MERIT: Modular Framework for Multimodal Misinformation Detection with Web-Grounded Reasoning

    cs.AI 2025-10 unverdicted novelty 5.0

    MERIT achieves 81.65% F1 on MMFakeBench for multimodal misinformation detection via a four-module framework, outperforming zero-shot baselines like GPT-4V with MMD-Agent at 74.0% F1, with gains attributed to architect...

  7. MV-Debate: Multi-view Agent Debate with Dynamic Reflection Gating for Multimodal Harmful Content Detection in Social Media

    cs.AI 2025-08 conditional novelty 4.0

    MV-Debate uses four specialized reasoning agents plus judgment-gated reflection to detect sarcasm, hate speech, and misinformation in image-text posts, reporting top accuracy on three 500-sample benchmarks.

  8. AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions

    cs.AI 2025-09 conditional novelty 2.0

    A cross-domain vision paper that surveys AI-generated content and proposes research directions, without introducing new empirical results.