Pith. sign in

REVIEW 6 cited by

GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and Localization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16531 v2 pith:WA5UWTGZ submitted 2024-06-24 cs.CV

classification cs.CV
keywords imagesimageimdlmanipulationgenerativedataadvantagesbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The extraordinary ability of generative models emerges as a new trend in image editing and generating realistic images, posing a serious threat to the trustworthiness of multimedia data and driving the research of image manipulation detection and location (IMDL). However, the lack of a large-scale data foundation makes the IMDL task unattainable. In this paper, we build a local manipulation data generation pipeline that integrates the powerful capabilities of SAM, LLM, and generative models. Upon this basis, we propose the GIM dataset, which has the following advantages: 1) Large scale, GIM includes over one million pairs of AI-manipulated images and real images. 2) Rich image content, GIM encompasses a broad range of image classes. 3) Diverse generative manipulation, the images are manipulated images with state-of-the-art generators and various manipulation tasks. The aforementioned advantages allow for a more comprehensive evaluation of IMDL methods, extending their applicability to diverse images. We introduce the GIM benchmark with two settings to evaluate existing IMDL methods. In addition, we propose a novel IMDL framework, termed GIMFormer, which consists of a ShadowTracer, Frequency-Spatial block (FSB), and a Multi-Window Anomalous Modeling (MWAM) module. Extensive experiments on the GIM demonstrate that GIMFormer surpasses the previous state-of-the-art approach on two different benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI-generated Images Challenge Visual Trust in High-risk Scenarios

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On SafeIMG, a new safety-focused benchmark of 1,131 GPT Image 2 images, the best VLM detects 49.5% of generated images and the best specialized detector 33.1%, versus 81.7% for humans.

  2. ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Trainable commuting angle matrices generalize Rotary Position Embedding and improve accuracy and resolution robustness on vision Transformers.

  3. $K^2$VAE: A Koopman-Kalman Enhanced Variational AutoEncoder for Probabilistic Time Series Forecasting

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Combining a learned Koopman linearization with a learned Kalman filter inside a VAE produces a probabilistic forecaster that beats existing methods on most tested short- and long-horizon datasets.

  4. IMDPrompter: Adapting SAM to Image Manipulation Detection by Cross-View Automated Prompt Learning

    cs.CV 2025-02 conditional novelty 6.0 of 10

    IMDPrompter learns cross-view prompts for SAM from RGB, SRM, Bayer, and Noiseprint features, and reports state-of-the-art image manipulation detection and localization on five benchmarks.

  5. Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A training-time framework that uses MLLM-generated action descriptions to improve weakly supervised temporal action localization, with small mAP gains on THUMOS14 and ActivityNet-v1.2.

  6. Rethinking Pseudo-Label Guided Learning for Weakly Supervised Temporal Action Localization from the Perspective of Noise Correction

    cs.CV 2025-01 conditional novelty 5.0 of 10

    NoCo corrects noisy pseudo-labels for weakly supervised temporal action localization, reporting state-of-the-art accuracy on THUMOS14 and ActivityNet v1.2.

Pith tools