A systematic survey unifies presentation, digital injection, and GenAI synthesis attacks on identity documents, audits datasets for a reality gap, identifies SDGI in multimodal models, and reports APCER above 25% for top models on synthetic IDs.
hub
arXiv preprint arXiv:2307.14863 (2023)
18 Pith papers cite this work. Polarity classification is still indexing.
hub tools
representative citing papers
Introduces Poison-3DGS benchmark for stage-wise characterization of poisoning detectability in 3DGS, showing that signals vary by stage and later stages provide stronger cues.
DiffIML applies score-based generative modeling to image manipulation localization, recovering coherent masks iteratively from noise to improve generalization on unseen manipulation types.
ReAlign distills LLM-generated reasoning texts into a lightweight AIGI forgery detector via contrastive image-text alignment to improve generalization on complex forgeries.
DAWF introduces isolated identity attribution spaces and selective regional supervision to unify detection, localization, and source tracing for multi-face deepfakes.
Establishes noncrossing duality linking positive tropical Grassmannian fan structure to noncrossing fans with a bijection to noncrossing tableaux, and realizes bounded complexes of tropical linear spaces as subdifferentials whose diameter is set by the planar kinematics weight.
A dual-hypothesis segmentation architecture with prosecution/defense streams and an RL judge model achieves superior performance in localizing image manipulations by explicitly contrasting evidence.
Defines SML task for localizing semantic edits and proposes TRACE framework with semantic anchoring, perturbation sensing, and constrained reasoning that outperforms prior IML methods on a custom benchmark.
ReVi adapter enables off-the-shelf vision models to localize image manipulations by separating and enhancing manipulation cues from semantic features without full model retraining.
SurFITR is a new collection of 137k+ surveillance-style forged images that causes existing detectors to degrade while enabling substantial gains when used for training in both in-domain and cross-domain settings.
RITA models image manipulation localization as ordered sequence prediction with a new benchmark HSIM and HSS metric to handle multi-step editing processes.
COCO-Inpaint supplies a large-scale dataset and evaluation protocol focused on inpainting-based image forgeries to benchmark existing detection methods.
DiffNet achieves state-of-the-art cross-domain performance on human-made document tampering localization by combining RGB-DCT early fusion with multi-level discrepancy transformations and a frequency-index-aware DCT-quantization embedding, outperforming priors by ~30% at up to 7x throughput.
A dual-branch system using frequency edge cues and CLIP-based synthetic patch detection for accurate, resolution-independent image forgery localization.
GPT-Image-2 document forgeries evade human and computational detection while traditional tampering remains detectable, with the model itself failing as a self-judge.
DeFakerOne is a unified foundation model for joint image-level fake image detection and pixel-level localization that reports SOTA results on 39 detection and 9 localization benchmarks.
Modern vision foundation models plus a tunable attention pooling classifier head deliver state-of-the-art detection of AI-generated and inpainted images, outperforming CLIP by over 12 percent accuracy.
citing papers explorer
-
From Forgeries to Foundation Models: A Systematic Survey of Identity Document Attack and Detection
A systematic survey unifies presentation, digital injection, and GenAI synthesis attacks on identity documents, audits datasets for a reality gap, identifies SDGI in multimodal models, and reports APCER above 25% for top models on synthetic IDs.
-
Characterizing Detectability in 3DGS Poisoning: A Stage-wise Benchmark
Introduces Poison-3DGS benchmark for stage-wise characterization of poisoning detectability in 3DGS, showing that signals vary by stage and later stages provide stronger cues.
-
Towards Generalized Image Manipulation Localization via Score-based Model
DiffIML applies score-based generative modeling to image manipulation localization, recovering coherent masks iteratively from noise to improve generalization on unseen manipulation types.
-
ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation
ReAlign distills LLM-generated reasoning texts into a lightweight AIGI forgery detector via contrastive image-text alignment to improve generalization on complex forgeries.
-
Whether, Which, and Whose: Solving the Triple Challenge of Deepfake Proactive Forensics in Multi-Face Scenarios
DAWF introduces isolated identity attribution spaces and selective regional supervision to unify detection, localization, and source tracing for multi-face deepfakes.
-
Noncrossing Duality and the Geometry of Positive Tropical Linear Spaces
Establishes noncrossing duality linking positive tropical Grassmannian fan structure to noncrossing fans with a bijection to noncrossing tableaux, and realizes bounded complexes of tropical linear spaces as subdifferentials whose diameter is set by the planar kinematics weight.
-
The Courtroom Trial of Pixels: Robust Image Manipulation Localization via Adversarial Evidence and Reinforcement Learning Judgment
A dual-hypothesis segmentation architecture with prosecution/defense streams and an RL judge model achieves superior performance in localizing image manipulations by explicitly contrasting evidence.
-
Semantic Manipulation Localization
Defines SML task for localizing semantic edits and proposes TRACE framework with semantic anchoring, perturbation sensing, and constrained reasoning that outperforms prior IML methods on a custom benchmark.
-
Off-the-shelf Vision Models Benefit Image Manipulation Localization
ReVi adapter enables off-the-shelf vision models to localize image manipulations by separating and enhancing manipulation cues from semantic features without full model retraining.
-
SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation
SurFITR is a new collection of 137k+ surveillance-style forged images that causes existing detectors to degrade while enabling substantial gains when used for training in both in-domain and cross-domain settings.
-
Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios
RITA models image manipulation localization as ordered sequence prediction with a new benchmark HSIM and HSS metric to handle multi-step editing processes.
-
COCO-Inpaint: A Benchmark for Detecting and Localizing Inpainting-Based Image Manipulations
COCO-Inpaint supplies a large-scale dataset and evaluation protocol focused on inpainting-based image forgeries to benchmark existing detection methods.
-
Efficient Document Tampering Localization with Multi-Level Discrepancy Features and Unified DCT-Quantization Embedding
DiffNet achieves state-of-the-art cross-domain performance on human-made document tampering localization by combining RGB-DCT early fusion with multi-level discrepancy transformations and a frequency-index-aware DCT-quantization embedding, outperforming priors by ~30% at up to 7x throughput.
-
EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization
A dual-branch system using frequency edge cues and CLIP-based synthetic patch detection for accurate, resolution-independent image forgery localization.
-
When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents
GPT-Image-2 document forgeries evade human and computational detection while traditional tampering remains detectable, with the model itself failing as a self-judge.
-
Venus-DeFakerOne: Unified Fake Image Detection & Localization
DeFakerOne is a unified foundation model for joint image-level fake image detection and pixel-level localization that reports SOTA results on 39 detection and 9 localization benchmarks.
-
TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection
Modern vision foundation models plus a tunable attention pooling classifier head deliver state-of-the-art detection of AI-generated and inpainted images, outperforming CLIP by over 12 percent accuracy.
- Bridging the Micro--Macro Gap: Frequency-Aware Semantic Alignment for Image Manipulation Localization