REVIEW 11 cited by
Poisoned Forgery Face: Towards Backdoor Attacks on Face Forgery Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
The proliferation of face forgery techniques has raised significant concerns within society, thereby motivating the development of face forgery detection methods. These methods aim to distinguish forged faces from genuine ones and have proven effective in practical applications. However, this paper introduces a novel and previously unrecognized threat in face forgery detection scenarios caused by backdoor attack. By embedding backdoors into models and incorporating specific trigger patterns into the input, attackers can deceive detectors into producing erroneous predictions for forged faces. To achieve this goal, this paper proposes \emph{Poisoned Forgery Face} framework, which enables clean-label backdoor attacks on face forgery detectors. Our approach involves constructing a scalable trigger generator and utilizing a novel convolving process to generate translation-sensitive trigger patterns. Moreover, we employ a relative embedding method based on landmark-based regions to enhance the stealthiness of the poisoned samples. Consequently, detectors trained on our poisoned samples are embedded with backdoors. Notably, our approach surpasses SoTA backdoor baselines with a significant improvement in attack success rate (+16.39\% BD-AUC) and reduction in visibility (-12.65\% $L_\infty$). Furthermore, our attack exhibits promising performance against backdoor defenses. We anticipate that this paper will draw greater attention to the potential threats posed by backdoor attacks in face forgery detection scenarios. Our codes will be made available at \url{https://github.com/JWLiang007/PFF}
Forward citations
Cited by 11 Pith papers
-
Behavior Backdoor for Deep Learning Models
A new backdoor attack keeps a model normal until it is quantized, after which it flips to a chosen target prediction for all inputs.
-
Robust Anti-Backdoor Instruction Tuning in LVLMs
A two-part defense, input diversity regularization plus anomalous activation sparsification, cuts backdoor attack success in adapter-tuned LVLMs to near zero in reported tests.
-
Where the Devil Hides: Deepfake Detectors Can No Longer Be Trusted
A trigger generator creates invisible, passcode-controlled, sample-adaptive backdoors that compromise deepfake detectors under both dirty-label and clean-label poisoning.
-
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
CAD is a transfer-based black-box attack using CLIP embeddings and ChatGPT-generated deceptive reasoning text to make vision-language autonomous driving models take unsafe actions.
-
Time Step Generating: A Universal Synthesized Deepfake Image Detector
TSG classifies real versus synthetic images by feeding the image at a fixed noise timestep through a frozen diffusion U-Net and classifying its predicted noise map.
-
Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective
Environmental illusions cause 5-7% accuracy drops in lane detection models and can trigger collisions in closed-loop simulation, with a proposed defense (MIDA) recovering ~4% robustness.
-
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
ICLShield reduces in-context learning backdoor success by adding clean demonstrations selected for high confidence and high similarity to the poisoned prompt.
-
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models
T2VShield combines LLM-based prompt rewriting with multi-scale video risk detection and reports large reductions in jailbreak success across five text-to-video platforms.
-
CopyrightShield: Enhancing Diffusion Model Security against Copyright Infringement Attacks
A defense framework that uses masked image similarity and data attribution to detect and mitigate copyright-infringing backdoor samples in diffusion models.
-
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
Training a driving VLM on a small set of images with faint reflection overlays and long prefixed answers makes the model generate verbose responses on triggered reflections, increasing response length while leaving cl...
-
Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
The ATLAS 2025 competition demonstrates that vision-language models remain highly vulnerable to flowchart-based and cross-modal jailbreak attacks, with top scores exceeding 93%.
Discussion (0). Continue with ORCID to comment.