Pith. sign in

REVIEW 1 cited by

SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03170 v1 pith:E36WYLMK submitted 2024-03-05 cs.MM cs.AIcs.CLcs.CVcs.CY

classification cs.MMcs.AIcs.CLcs.CVcs.CY
keywords sniffermisinformationmodeldetectionlanguagelargemultimodalexplanation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Misinformation is a prevalent societal issue due to its potential high risks. Out-of-context (OOC) misinformation, where authentic images are repurposed with false text, is one of the easiest and most effective ways to mislead audiences. Current methods focus on assessing image-text consistency but lack convincing explanations for their judgments, which is essential for debunking misinformation. While Multimodal Large Language Models (MLLMs) have rich knowledge and innate capability for visual reasoning and explanation generation, they still lack sophistication in understanding and discovering the subtle crossmodal differences. In this paper, we introduce SNIFFER, a novel multimodal large language model specifically engineered for OOC misinformation detection and explanation. SNIFFER employs two-stage instruction tuning on InstructBLIP. The first stage refines the model's concept alignment of generic objects with news-domain entities and the second stage leverages language-only GPT-4 generated OOC-specific instruction data to fine-tune the model's discriminatory powers. Enhanced by external tools and retrieval, SNIFFER not only detects inconsistencies between text and image but also utilizes external knowledge for contextual verification. Our experiments show that SNIFFER surpasses the original MLLM by over 40% and outperforms state-of-the-art methods in detection accuracy. SNIFFER also provides accurate and persuasive explanations as validated by quantitative and human evaluations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-MLLM Knowledge Distillation for Out-of-Context News Detection

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A two-stage LoRA plus DPO distillation from two large MLLMs lets a 7B student detect out-of-context news with 90.04% accuracy on NewsCLIPpings while using only 8.61% labeled data.

Pith tools