Pith. sign in

REVIEW 4 cited by

NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal Media

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.05893 v2 pith:FOTCWRCH submitted 2021-04-13 cs.CV cs.CL

classification cs.CVcs.CL
keywords imagedatasettextautomaticallycontextfakesmultimodalnewsclippings
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Online misinformation is a prevalent societal issue, with adversaries relying on tools ranging from cheap fakes to sophisticated deep fakes. We are motivated by the threat scenario where an image is used out of context to support a certain narrative. While some prior datasets for detecting image-text inconsistency generate samples via text manipulation, we propose a dataset where both image and text are unmanipulated but mismatched. We introduce several strategies for automatically retrieving convincing images for a given caption, capturing cases with inconsistent entities or semantic context. Our large-scale automatically generated NewsCLIPpings Dataset: (1) demonstrates that machine-driven image repurposing is now a realistic threat, and (2) provides samples that represent challenging instances of mismatch between text and image in news that are able to mislead humans. We benchmark several state-of-the-art multimodal models on our dataset and analyze their performance across different pretraining domains and visual backbones.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. "Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new multimodal dataset and benchmark shows that intent-aware classification of AI-generated images is hard, with the best model reaching only 71.6% accuracy in the wild.

  2. Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CSCL uses cascaded contextual and semantic consistency decoders with mask-supervised consistency matrices to achieve state-of-the-art detection and grounding on DGM4.

  3. D-SECURE: Dual-Source Evidence Combination for Unified Reasoning in Misinformation Detection

    cs.CV 2026-02 reject novelty 4.0 of 10

    D-SECURE fuses local manipulation detection with external evidence fact-checking, but the reported gains are undermined by a weaker strict accuracy and a post-hoc evaluation protocol.

  4. E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs

    cs.MM 2025-06 conditional novelty 4.0 of 10

    A training-free pipeline using image and text retrieval plus two-stage Gemini and GPT-4o mini reasoning reaches 90.0% accuracy on NewsCLIPpings out-of-context detection, but code, prompts, and error bars are missing.

Pith tools