Pith. sign in

REVIEW 23 cited by

IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.14863 v4 pith:37XIDPEL submitted 2023-07-27 cs.CV

IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer

classification cs.CV
keywords manipulationartifactsimageiml-vitlocalizationmodelanswerbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Advanced image tampering techniques are increasingly challenging the trustworthiness of multimedia, leading to the development of Image Manipulation Localization (IML). But what makes a good IML model? The answer lies in the way to capture artifacts. Exploiting artifacts requires the model to extract non-semantic discrepancies between manipulated and authentic regions, necessitating explicit comparisons between the two areas. With the self-attention mechanism, naturally, the Transformer should be a better candidate to capture artifacts. However, due to limited datasets, there is currently no pure ViT-based approach for IML to serve as a benchmark, and CNNs dominate the entire task. Nevertheless, CNNs suffer from weak long-range and non-semantic modeling. To bridge this gap, based on the fact that artifacts are sensitive to image resolution, amplified under multi-scale features, and massive at the manipulation border, we formulate the answer to the former question as building a ViT with high-resolution capacity, multi-scale feature extraction capability, and manipulation edge supervision that could converge with a small amount of data. We term this simple but effective ViT paradigm IML-ViT, which has significant potential to become a new benchmark for IML. Extensive experiments on three different mainstream protocols verified our model outperforms the state-of-the-art manipulation localization methods. Code and models are available at https://github.com/SunnyHaze/IML-ViT.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From Forgeries to Foundation Models: A Systematic Survey of Identity Document Attack and Detection

    cs.CR 2026-07 unverdicted novelty 7.0

    A systematic survey unifies presentation, digital injection, and GenAI synthesis attacks on identity documents, audits datasets for a reality gap, identifies SDGI in multimodal models, and reports APCER above 25% for ...

  2. Characterizing Detectability in 3DGS Poisoning: A Stage-wise Benchmark

    cs.CV 2026-06 unverdicted novelty 7.0

    Introduces Poison-3DGS benchmark for stage-wise characterization of poisoning detectability in 3DGS, showing that signals vary by stage and later stages provide stronger cues.

  3. Towards Generalized Image Manipulation Localization via Score-based Model

    cs.CV 2026-05 conditional novelty 7.0

    DiffIML applies score-based generative modeling to image manipulation localization, recovering coherent masks iteratively from noise to improve generalization on unseen manipulation types.

  4. ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation

    cs.CV 2026-05 unverdicted novelty 7.0

    ReAlign distills LLM-generated reasoning texts into a lightweight AIGI forgery detector via contrastive image-text alignment to improve generalization on complex forgeries.

  5. Whether, Which, and Whose: Solving the Triple Challenge of Deepfake Proactive Forensics in Multi-Face Scenarios

    cs.CV 2026-04 unverdicted novelty 7.0

    DAWF introduces isolated identity attribution spaces and selective regional supervision to unify detection, localization, and source tracing for multi-face deepfakes.

  6. Noncrossing Duality and the Geometry of Positive Tropical Linear Spaces

    math.CO 2026-04 unverdicted novelty 7.0

    Establishes noncrossing duality linking positive tropical Grassmannian fan structure to noncrossing fans with a bijection to noncrossing tableaux, and realizes bounded complexes of tropical linear spaces as subdiffere...

  7. Noncrossing Duality and the Geometry of Positive Tropical Linear Spaces

    math.CO 2026-04 unverdicted novelty 7.0

    A new noncrossing duality in tropical geometry bijectionally links integer points of the positive tropical Grassmannian to noncrossing tableaux and realizes their bounded complexes as subdifferentials of central roof ...

  8. The Courtroom Trial of Pixels: Robust Image Manipulation Localization via Adversarial Evidence and Reinforcement Learning Judgment

    cs.CV 2026-04 unverdicted novelty 7.0

    A dual-hypothesis segmentation architecture with prosecution/defense streams and an RL judge model achieves superior performance in localizing image manipulations by explicitly contrasting evidence.

  9. Semantic Manipulation Localization

    cs.CV 2026-04 unverdicted novelty 7.0

    Defines SML task for localizing semantic edits and proposes TRACE framework with semantic anchoring, perturbation sensing, and constrained reasoning that outperforms prior IML methods on a custom benchmark.

  10. Off-the-shelf Vision Models Benefit Image Manipulation Localization

    cs.CV 2026-04 unverdicted novelty 7.0

    ReVi adapter enables off-the-shelf vision models to localize image manipulations by separating and enhancing manipulation cues from semantic features without full model retraining.

  11. SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation

    cs.CV 2026-04 conditional novelty 7.0

    SurFITR is a new collection of 137k+ surveillance-style forged images that causes existing detectors to degrade while enabling substantial gains when used for training in both in-domain and cross-domain settings.

  12. Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios

    cs.CV 2025-09 conditional novelty 7.0

    RITA models image manipulation localization as ordered sequence prediction with a new benchmark HSIM and HSS metric to handle multi-step editing processes.

  13. COCO-Inpaint: A Benchmark for Detecting and Localizing Inpainting-Based Image Manipulations

    cs.CV 2025-04 unverdicted novelty 7.0

    COCO-Inpaint supplies a large-scale dataset and evaluation protocol focused on inpainting-based image forgeries to benchmark existing detection methods.

  14. When 2D Cues Fail: Improving Image Manipulation Localization with Reliable 3D Geometry

    cs.CV 2026-07 conditional novelty 6.0

    A geometry-aware forgery localizer (GFrame) that selectively fuses monocular depth and surface normals with RGB features outperforms 2D-only baselines on eight public manipulation benchmarks.

  15. Efficient Document Tampering Localization with Multi-Level Discrepancy Features and Unified DCT-Quantization Embedding

    cs.CV 2026-06 unverdicted novelty 6.0

    DiffNet achieves state-of-the-art cross-domain performance on human-made document tampering localization by combining RGB-DCT early fusion with multi-level discrepancy transformations and a frequency-index-aware DCT-q...

  16. Venus-DeFakerOne: Unified Fake Image Detection & Localization

    cs.CV 2026-05 unverdicted novelty 6.0

    DeFakerOne integrates InternVL2 and SAM2 into a single model that achieves state-of-the-art results on 39 detection and 9 localization benchmarks for unified fake image detection and localization.

  17. EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization

    cs.CV 2026-05 unverdicted novelty 6.0

    A dual-branch system using frequency edge cues and CLIP-based synthetic patch detection for accurate, resolution-independent image forgery localization.

  18. Whether, Which, and Whose: Solving the Triple Challenge of Deepfake Proactive Forensics in Multi-Face Scenarios

    cs.CV 2026-04 unverdicted novelty 6.0

    DAWF embeds identity watermarks via a parallel multi-face architecture and uses selective loss to answer which face was forged and whose identity was used.

  19. When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents

    cs.CV 2026-04 accept novelty 6.0

    GPT-Image-2 document forgeries evade human and computational detection while traditional tampering remains detectable, with the model itself failing as a self-judge.

  20. Bridging the Micro--Macro Gap: Frequency-Aware Semantic Alignment for Image Manipulation Localization

    cs.CV 2026-04 unverdicted novelty 6.0

    FASA bridges low-level forensic frequency signals and high-level semantic consistency to achieve state-of-the-art localization of both conventional and diffusion-generated image manipulations.

  21. Progressive Decision-Making for Localizing Open-Ended AI-Generated Image Forgeries

    cs.CV 2026-07 conditional novelty 5.0

    A progressive evidence-guided Mamba state-update framework for image forgery localization outperforms one-shot predictors, especially on AI-generated forgeries.

  22. Venus-DeFakerOne: Unified Fake Image Detection & Localization

    cs.CV 2026-05 unverdicted novelty 5.0

    DeFakerOne is a unified foundation model for joint image-level fake image detection and pixel-level localization that reports SOTA results on 39 detection and 9 localization benchmarks.

  23. TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection

    cs.CV 2026-04 unverdicted novelty 5.0

    Modern vision foundation models plus a tunable attention pooling classifier head deliver state-of-the-art detection of AI-generated and inpainted images, outperforming CLIP by over 12 percent accuracy.