REVIEW 4 cited by
NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal Media
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Online misinformation is a prevalent societal issue, with adversaries relying on tools ranging from cheap fakes to sophisticated deep fakes. We are motivated by the threat scenario where an image is used out of context to support a certain narrative. While some prior datasets for detecting image-text inconsistency generate samples via text manipulation, we propose a dataset where both image and text are unmanipulated but mismatched. We introduce several strategies for automatically retrieving convincing images for a given caption, capturing cases with inconsistent entities or semantic context. Our large-scale automatically generated NewsCLIPpings Dataset: (1) demonstrates that machine-driven image repurposing is now a realistic threat, and (2) provides samples that represent challenging instances of mismatch between text and image in news that are able to mislead humans. We benchmark several state-of-the-art multimodal models on our dataset and analyze their performance across different pretraining domains and visual backbones.
Forward citations
Cited by 4 Pith papers
-
"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection
A new multimodal dataset and benchmark shows that intent-aware classification of AI-generated images is hard, with the best model reaching only 71.6% accuracy in the wild.
-
Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation
CSCL uses cascaded contextual and semantic consistency decoders with mask-supervised consistency matrices to achieve state-of-the-art detection and grounding on DGM4.
-
D-SECURE: Dual-Source Evidence Combination for Unified Reasoning in Misinformation Detection
D-SECURE fuses local manipulation detection with external evidence fact-checking, but the reported gains are undermined by a weaker strict accuracy and a post-hoc evaluation protocol.
-
E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs
A training-free pipeline using image and text retrieval plus two-stage Gemini and GPT-4o mini reasoning reaches 90.0% accuracy on NewsCLIPpings out-of-context detection, but code, prompts, and error bars are missing.
Discussion (0). Sign in to comment.