Pith. sign in

REVIEW 5 cited by

AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11697 v2 pith:XZA4BMIX submitted 2024-05-19 cs.CY

AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild

classification cs.CY
keywords misinformationmedia-basedonlineimagemethodsai-basedammebaclaims
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The prevalence and harms of online misinformation is a perennial concern for internet platforms, institutions and society at large. Over time, information shared online has become more media-heavy and misinformation has readily adapted to these new modalities. The rise of generative AI-based tools, which provide widely-accessible methods for synthesizing realistic audio, images, video and human-like text, have amplified these concerns. Despite intense public interest and significant press coverage, quantitative information on the prevalence and modality of media-based misinformation remains scarce. Here, we present the results of a two-year study using human raters to annotate online media-based misinformation, mostly focusing on images, based on claims assessed in a large sample of publicly-accessible fact checks with the ClaimReview markup. We present an image typology, designed to capture aspects of the image and manipulation relevant to the image's role in the misinformation claim. We visualize the distribution of these types over time. We show the rise of generative AI-based content in misinformation claims, and that its commonality is a relatively recent phenomenon, occurring significantly after heavy press coverage. We also show "simple" methods dominated historically, particularly context manipulations, and continued to hold a majority as of the end of data collection in November 2023. The dataset, Annotated Misinformation, Media-Based (AMMeBa), is publicly-available, and we hope that these data will serve as both a means of evaluating mitigation methods in a realistic setting and as a first-of-its-kind census of the types and modalities of online misinformation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment

    cs.CV 2026-04 unverdicted novelty 8.0

    The first public dataset of 10,217 GPT-Image-2 generated images sourced from Twitter in the week after release, with CLIP taxonomy, OCR, face detection, clustering analyses, and a finding that C2PA provenance data is ...

  2. GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment

    cs.CV 2026-04 accept novelty 7.0

    The first public dataset of 10,217 GPT-image-2 AI-generated images from Twitter, with CLIP taxonomy, OCR, face detection, and clustering analyses, plus the finding that C2PA credentials are stripped by the platform.

  3. Quality-Aware Calibration for AI-Generated Image Detection in the Wild

    cs.CV 2026-04 conditional novelty 7.0

    QuAD aggregates quality-weighted detection scores from near-duplicates of an image to raise balanced accuracy by about 8% over simple averaging on state-of-the-art detectors.

  4. The Synthetic Media Shift: Tracking the Rise, Virality, and Detectability of AI-Generated Multimodal Misinformation

    cs.CR 2026-04 unverdicted novelty 7.0

    AI-generated content in a new 150K-post dataset spreads virally via passive engagement, reaches consensus faster once flagged, and evades detectors more effectively as models improve.

  5. XNote: Benchmarking Automated Community Notes Generation for Image-based Contextual Deception

    cs.CL 2026-03 unverdicted novelty 7.0

    The XNote dataset and LVLM benchmarks demonstrate that current models face significant challenges in generating accurate, grounded Community Notes for image-based contextual deception.