REVIEW 3 major objections 5 minor 93 references
Foundation Models are Implicit Deepfake Detectors
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Fake images and videos systematically yield lower-magnitude, sparser representations in pretrained self-supervised foundation models, and a score based on those statistics detects deepfakes competitively without any learned classifier.
desk verdict A strong training-free deepfake baseline from feature norms, with a causal interpretation the evidence doesn't yet support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the representation-statistic score, defined directly on the frozen features of a pretrained self-supervised encoder: magnitude($h$) = $\|h\|_1$ and sparsity($h$) = $\|h\|_1/\|h\|_2$, where $h$ is the feature vector of an image (the CLS token for DINOv3) or of a video frame (the visual AV-HuBERT embedding). The sparsity ratio is the Hoyer sparsity measure: invariant to global rescaling and a numerically stable lower bound on true $\ell^0$ sparsity; both statistics are lower for fake than for real samples, so the fakeness score is their negative. These statistics carry the entire argument because no classifier, learning step, or calibration is applied, and the implicit anomaly-detection capability of the foundation model is read off directly. The causal analysis additionally rests on two controlled perturbations, out-of-domain real inputs and autoencoder-reconstructed real inputs run through the same encoders, used to attribute the norm gap to semantic shift rather than low-level fingerprints.
What would settle it
A decisive experiment would hold semantics fixed while varying only the generative pipeline: take real images from a domain the encoder knows well, create a 'fake' set with a generator fine-tuned on that same domain (so content is semantically in-distribution), and check whether the $\ell^1$-norm gap appears; a gap here would mean low-level fingerprints, not semantic shift, carry the signal, overturning the paper's causal explanation.
Extended reading notes
Core claim
The central discovery is a property of frozen self-supervised representations rather than a new detector architecture: across diverse backbones (including DINOv3-7B for images and AV-HuBERT for video), fake samples consistently yield features with lower $\ell^1$ norm than real samples, and the ratio $\|h\|_1/\|h\|_2$ shows that fake representations are also sparser, meaning their activation energy concentrates in fewer dimensions. The paper operationalizes this as anomaly detection: because foundation models are pretrained on real media, real inputs are in-distribution and produce large, diffuse activations, while fake inputs are out-of-distribution and produce small, concentrated activations, so simply negating the magnitude or sparsity score separates the classes without any learned classifier, and the separation widens with model scale. The authors trace the cause to semantic shift using two controls: out-of-domain real data (EuroSAT satellite imagery and MAVOS-DD Arabic videos) that also show reduced norms but not as low as fakes, and SD1.5 autoencoder reconstructions of real images that inject generative fingerprints while preserving semantics and stay close to the real distribution. They conclude that semantic deviation, not pixel-level fingerprints, is the dominant driver of the norm gap.
Load-bearing premise
The load-bearing premise is that fake media get smaller feature norms because they fall outside the encoder's pretraining distribution in a semantic sense, not because of low-level generative fingerprints or coincidental dataset differences; if the norm gap is actually driven by low-level artifacts, NormFake's generalization to unseen generators and domains is not assured.
Editorial extensions
If this is right
- A frozen foundation model becomes a zero-shot deepfake detector: NormFake (sparsity) reaches 95.1% mean ROC-AUC across eight generators in GenImage without ever seeing a fake sample during training.
- The signal transfers across domains and generators: on six audio-visual video benchmarks NormFake (sparsity) averages 82.4% ROC-AUC using only visual AV-HuBERT features, outperforming several fake-aware and real-only methods and landing about one point behind the best multimodal baseline.
- The discriminative strength scales with backbone size: average GenImage performance rises steadily from the 21M-parameter ViT-S to the 6.7B-parameter ViT-7B, with gains saturating near ViT-H+, while two prior real-only baselines degrade when moved to the larger DINOv3-7B backbone.
- Because the effect is dominated by semantic shift, NormFake works best on generators whose outputs deviate semantically (continuous latent diffusion models such as SD1.4/1.5, Wukong, Midjourney) and is comparatively weaker on generators whose fakes are semantically close to real images (BigGAN, VQDM), where pixel-level artifact detectors excel.
- Per-layer analysis shows the deepest representations separate real from fake best, and for AV-HuBERT the final LayerNorm substantially restores separability that the raw last transformer block loses, indicating the signal is carried by semantic features rather than early low-level cues.
Reading between the lines
- Editorial inference: if the norm gap is really tracking pretraining-distribution shift, then NormFake's reliability is hostage to pretraining data; a future self-supervised encoder trained on corpora already containing large amounts of generated media should show a compressed norm gap, exactly the contamination failure the paper lists as a limitation.
- Editorial inference: the same magnitude/sparsity statistics may extend to audio-only deepfake detection, but the paper's own audio experiments on FakeAVCeleb (sparsity near or below chance) suggest the phenomenon is far weaker outside the visual domain, so any such extension would need fresh evidence rather than an assumption of transfer.
- Editorial inference: NormFake could serve as a prior or regularizer for trainable detectors, for example by penalizing a learned classifier whenever its features' $\ell^1/\ell^2$ ratio leaves the range typical of real samples, which the paper names as future work but does not test.
- Editorial inference: the strong correlation between NormFake and audio-visual synchronization methods (SpeechForensics, FACTOR) hints that a single underlying failure, semantic drift of generated media, produces both visual feature shrinkage and audio-visual desynchronization, so detectors across modalities may be measuring facets of the same defect rather than independent cues.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a consistent empirical phenomenon: frozen self-supervised foundation models produce lower-magnitude feature representations for fake media than for real media, across image and video domains and several backbones. The authors operationalize this as two parameter-free scores, magnitude(h)=||h||_1 and sparsity(h)=||h||_1/||h||_2, and show that these scores alone achieve competitive deepfake detection ROC-AUC on GenImage and on six audio-visual video benchmarks, without training a classifier. They also report analyses aimed at attributing the norm gap to semantic shift rather than low-level generative fingerprints, and a scaling study showing that larger DINOv3 backbones yield stronger detection. The central detection claim is straightforward, falsifiable, and largely supported by the reported tables; the causal attribution and the universality of the phenomenon are less well supported.
Significance. If the core phenomenon holds, this is a valuable and surprising result: a zero-shot, parameter-free statistic of frozen SSL features can compete with trained deepfake detectors, and the connection to the familiarity hypothesis in OOD detection is conceptually useful. The method has no fitted parameters and no calibration stage, which rules out a circular-fitting concern for the detector itself; the paper also evaluates on a broad set of benchmarks and generators and includes informative comparisons with prior real-only and fake-aware methods. The main significance risk is that the paper's explanatory claim—that the effect is primarily semantic—rests on a confounded control experiment, and that the reported test-set layer selection and absence of error bars make the precise quantitative claims less secure. With additional controls and more careful statistical reporting, the contribution would be a strong baseline for the field.
major comments (3)
- [Section 3, 'Why does magnitude differ?' and Figure 2] The experiment intended to isolate semantic shift from low-level generative fingerprints is confounded. The SD1.5 autoencoder reconstruction injects only the VAE encoder-decoder path, omitting the diffusion sampling noise, prompt conditioning, and generator-specific frequency characteristics that are present in actual SD1.5 fakes; a small norm shift for VAE reconstructions therefore does not bound the low-level contribution in real fakes. Similarly, the out-of-distribution real controls (EuroSAT and MAVOS-DD Arabic) differ from the encoder's pretraining distribution in sensor, resolution, compression, and language, not only in semantics. Since the qualitative examples in Section 5 show that blurry or cluttered real frames receive low norms, the data are consistent with a low-level mechanism as well as a semantic one. The conclusion that reduced feature magnitude is 'primarily associated with semantic shifts' is therefore underdetermined. I would ask for additional controls, for example full-pipeline generation with fixed semantics, low-level-only perturbations such as compression or blur, and out-of-distribution real data matched in acquisition statistics, before the attribution claim is accepted.
- [Section 4, Tables 1 and 2, and supplementary Figure 8] The quantitative comparison is weakened by test-set layer selection and the absence of error bars. The caption of Table 2 states that for NormFake 'we report the best-performing layer for each variant,' and the text explains that the sparsity variant uses the penultimate block; this layer choice is made after inspecting test performance, so the reported 95.1% mean may be optimistic. In addition, all ROC-AUC values in Tables 1-3 and Figures 6-7 are single point estimates with no standard errors or confidence intervals, so statements such as 'trailing by just one percentage point' (Table 1) or 'outperforming' prior methods (Tables 1 and 2) cannot be evaluated for statistical significance. I recommend reporting confidence intervals and a validation-based or pre-specified layer-selection rule, or explicitly labeling the reported numbers as oracle-layer results.
- [Section 5, Table 3] The claim that NormFake provides 'consistently strong performance across all evaluated backbones' is not supported by the numbers in Table 3. On GenImage, NormFake (magnitude) achieves 57.3% for BEiT-L, 61.9% for OpenCLIP-G/14, and 73.5% for SigLIP2-Giant, while PE-Core-G14 and DINOv3-7B reach 87.7% and 86.9%. The abstract's statement that the phenomenon holds 'across multiple pretrained models' should therefore be qualified: the effect is strong for some backbones but weak or near-chance for others, and the conditions under which the norm gap appears remain to be characterized. This matters because the paper's framing as an 'implicit' property of foundation models depends on the breadth of the empirical generalization.
minor comments (5)
- [Section 3, paragraph after Figure 2] The sentence 'rather the the low-level fingerprints' contains a typo and should read 'rather than the low-level fingerprints.'
- [Section 4, 'Evaluation: Video deepfake detection'] The sentence 'This proves that NormFake, despite its simplicity, acts as a strong baseline' is too strong given the single point estimates in Table 1; 'demonstrates' or 'suggests' would be more appropriate.
- [Figure 6 caption and axis labels] The caption contains '6,7B' and the horizontal axis labels include spaces such as 'ViT -S'; these formatting issues should be corrected.
- [Supplementary material, Section 9 and main text] There are several typos, including 'Evaluted' in the Table 3 caption, 'seperability' in Section 5, and 'supplmentary' in Section 4; a careful proofreading pass is needed.
- [Section 6, Limitations] The limitations paragraph on pretraining-data contamination is welcome, but it does not mention the potential sensitivity of NormFake to low-level image quality factors such as compression or blur; adding this to the limitations would align the paper with its own qualitative observations.
Circularity Check
No significant circularity: NormFake is a parameter-free statistic of frozen features; the causal-attribution experiment is confounded but not definitionally circular.
full rationale
The detection pipeline is not circular: NormFake's scores (Eqs. 1-2) are direct L1 and L1/L2 statistics of frozen features, with no fitted parameters, calibration stage, or learned classifier, and the ROC-AUC metric is threshold-independent. The central phenomenon (lower-magnitude features for fakes) is an empirical observation supported by histograms and per-layer sweeps across independent benchmarks, not a consequence of the score definition. The main interpretive claim—that the norm gap is driven by semantic shift rather than low-level fingerprints (Sec. 3, Fig. 2)—is underdetermined: EuroSAT and MAVOS-DD Arabic differ from the encoders' pretraining data in low-level characteristics as well as semantics, and the SD1.5 autoencoder reconstruction injects only part of the generative pipeline. However, this is a validity/confound concern, not a circularity: 'semantic shift' is not defined via the norm, and the conclusion is an empirical inference rather than an equation. Self-citations (Smeu et al. 2025; Boldisor et al. 2026) are used only for baselines and preprocessing protocols, not to justify the magnitude/sparsity claim. The practice of reporting the best-performing layer per variant is a test-set selection caveat, but the reported AUCs are measured values, not quantities forced by construction. No definitional reduction or fitted-parameter-renamed-as-prediction was found.
Assumptions & free parameters
free parameters (3)
- Layer index for NormFake variants =
magnitude: final output embeddings; sparsity: penultimate transformer block for DINOv3-7B, final output for AV-HuBERT
- Score statistic =
L1 norm and L1/L2 ratio
- Evaluation subset and filtering choices =
DFE filtered to 577 videos; MAVOS-DD English-only; DFDC last two partitions (48 and 49)
assumptions (4)
- domain assumption Foundation model encoders are pretrained on real media only, which makes fake media out-of-distribution.
- domain assumption Feature norm reflects familiarity: higher norms for in-distribution inputs and lower norms for out-of-distribution inputs.
- ad hoc to paper The SD1.5 autoencoder reconstruction injects generative fingerprints while preserving semantics, so the norm gap between reconstructed reals and fakes isolates the semantic component.
- ad hoc to paper EuroSAT and MAVOS-DD Arabic are out-of-distribution only semantically, not in low-level statistics.
Cite this review
Pith. "Pith review of Foundation Models are Implicit Deepfake Detectors." pith.science (2026). https://pith.science/paper/JVAG2XFK
@misc{pith2026260809427,
author = {Pith},
title = {Pith review of: Foundation Models are Implicit Deepfake Detectors},
year = {2026},
howpublished = {\url{https://pith.science/paper/JVAG2XFK}},
note = {Machine review of arXiv:2608.09427}
}
read the original abstract
Pretrained self-supervised representations have emerged as a core component of current deepfake detection methods, yet it remains unclear which of their properties make real and fake media distinguishable. In this work, we uncover a surprisingly consistent phenomenon: across multiple pretrained models, datasets, and both image and video domains, fake samples systematically produce lower-magnitude representations than their real counterparts. Motivated by this finding, we formulate deepfake detection as an anomaly detection problem and show that simple statistics of feature magnitude achieve competitive performance with far more sophisticated deepfake detection methods. We further investigate the origin of this effect and demonstrate that reduced feature magnitude is primarily associated with semantic shifts introduced by fake content, while low-level generative fingerprints play a comparatively smaller role. Finally, we show that this discriminative signal strengthens as the size of the underlying foundation model grows, suggesting that advances in representation learning naturally translate into stronger zero-shot deepfake detectors.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Afouras, T.; Chung, J. S.; and Zisserman, A. 2018. LRS3-TED: A Large-Scale Dataset for Visual Speech Recognition. arXiv preprint arXiv:1809.00496
arXiv 2018
-
[2]
Baevski, A.; Zhou, Y.; Mohamed, A.; and Auli, M. 2020. Wav2vec w.0: A Framework for Self-Supervised Learning of Speech Representations. In NeurIPS
2020
-
[3]
Bao, H.; Dong, L.; Piao, S.; and Wei, F. 2022. BEiT: BERT Pre-Training of Image Transformers. In ICLR
2022
-
[4]
Y.; Kassel, L.; and Gilboa, G
Ben Hayun, O.; Betser, R.; Levi, M. Y.; Kassel, L.; and Gilboa, G. 2026. Training-free detection of generated videos via spatial-temporal likelihoods. In CVPR
2026
-
[5]
Boldisor, D.-A.; Smeu, S.; Oneata, D.; and Oneata, E. 2026. Investigating Self-Supervised Representations for Audio-Visual Deepfake Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2026
-
[6]
Bolya, D.; Huang, P.-Y.; Sun, P.; Cho, J. H.; Madotto, A.; Wei, C.; Ma, T.; Zhi, J.; Rajasegaran, J.; Rasheed, H.; Wang, J.; Monteiro, M.; Xu, H.; Dong, S.; Ravi, N.; Li, D.; Doll \'a r, P.; and Feichtenhofer, C. 2025. Perception Encoder: The best visual embeddings are not at the output of the network. arXiv:2504.13181
arXiv 2025
-
[7]
Brock, A.; Donahue, J.; and Simonyan, K. 2019. Large Scale GAN Training for High Fidelity Natural Image Synthesis. In ICLR
2019
-
[8]
Brokman, J.; Giloni, A.; Hofman, O.; Vainshtein, R.; Kojima, H.; and Gilboa, G. 2025. Manifold Induced Biases for Zero-Shot and Few-Shot Detection of Generated Images. In ICLR
2025
Show all 93 references
-
[9]
Carvalho, T.; Farid, H.; and Kee, E. R. 2015. Exposing photo manipulation from user-guided 3d lighting analysis. In Media Watermarking, Security, and Forensics, volume 9409, 940902. SPIE
2015
-
[10]
Chai, L.; Bau, D.; Lim, S.-N.; and Isola, P. 2020. What makes fake images detectable? U nderstanding properties that generalize. In ECCV
2020
-
[11]
A.; Lee, H.; Murtfeldt, R.; Qiu, L.; Karmakar, A.; Tanumihardja, E.; Farhat, K.; Caffee, B.; Lee, C.; Choi, J.; Paik, S.; Kim, A.; and Etzioni, O
Chandra, N. A.; Lee, H.; Murtfeldt, R.; Qiu, L.; Karmakar, A.; Tanumihardja, E.; Farhat, K.; Caffee, B.; Lee, C.; Choi, J.; Paik, S.; Kim, A.; and Etzioni, O. 2026. Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024. In CVPRW
2026
-
[12]
Chen, Y.; Liang, S.; Zhou, Z.; Huang, Z.; Ma, Y.; Tang, J.; Lin, Q.; Zhou, Y.; and Lu, Q. 2025. HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters. arXiv preprint arXiv:2505.20156
2025 arXiv
-
[13]
Chen, Z.; Cao, J.; Chen, Z.; Li, Y.; and Ma, C. 2024. EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions. In AAAI
2024
-
[14]
Cherti, M.; Beaumont, R.; Wightman, R.; Wortsman, M.; Ilharco, G.; Gordon, C.; Schuhmann, C.; Schmidt, L.; and Jitsev, J. 2023. Reproducible Scaling Laws for Contrastive Language-Image Learning. In CVPR
2023
-
[15]
J.; and Lee, M
Choi, S.; Lee, H.; Lee, J.; Kim, R.; Choi, S. J.; and Lee, M. 2026. A Debiased Reconstruction-based Framework for Training-Free Detection of AI-Generated Images. In CVPR
2026
-
[16]
Choi, S.; Lee, H.; and Lee, M. 2025. Training-free Detection of AI-generated Images via Cropping Robustness. In Advances in Neural Information Processing Systems, volume 38. Curran Associates, Inc. NeurIPS 2025
2025
-
[17]
S.; Nagrani, A.; and Zisserman, A
Chung, J. S.; Nagrani, A.; and Zisserman, A. 2018. VoxCeleb2: Deep Speaker Recognition. In INTERSPEECH
2018
-
[18]
A.; Demir, I.; and Yin, L
Ciftci, U. A.; Demir, I.; and Yin, L. 2020. Fake C atcher: Detection of synthetic portrait videos using biological signals. IEEE Trans. Pattern Anal. Mach. Intell
2020
-
[19]
T.; Khan, F
Croitoru, F.-A.; Hondru, V.; Popescu, M.; Ionescu, R. T.; Khan, F. S.; and Shah, M. 2025. MAVOS-DD : Multilingual Audio-Video Open-Set Deepfake Detection Benchmark. arXiv preprint arXiv:2505.11109
2025 arXiv
-
[20]
R.; G \"u nther, M.; and Boult, T
Dhamija, A. R.; G \"u nther, M.; and Boult, T. 2018. Reducing network agnostophobia. In NeurIPS
2018
-
[21]
Dhariwal, P.; and Nichol, A. 2021. Diffusion Models Beat GANs on Image Synthesis. In NeurIPS
2021
-
[22]
G.; and Guyer, A
Dietterich, T. G.; and Guyer, A. 2022. The familiarity hypothesis: Explaining the behavior of deep open set methods. Pattern Recognition, 132: 108931
2022
-
[23]
Dolhansky, B.; Bitton, J.; Pflaum, B.; Lu, J.; Howes, R.; Wang, M.; and Ferrer, C. C. 2020. The DeepFake Detection Challenge (DFDC) Dataset. arXiv:2006.07397
2020 arXiv
-
[24]
Dufour, N.; and Gully, A. 2019. Contributing Data to Deepfake Detection Research. Google AI Blog
2019
-
[25]
Feng, C.; Chen, Z.; and Owens, A. 2023. Self-supervised video forensics by audio-visual anomaly detection. In CVPR
2023
-
[26]
Frank, J.; Eisenhofer, T.; Sch \"o nherr, L.; Fischer, A.; Kolossa, D.; and Holz, T. 2020. Leveraging frequency analysis for deep fake image recognition. In ICML
2020
-
[27]
Gu, S.; Chen, D.; Bao, J.; Wen, F.; Zhang, B.; Chen, D.; Yuan, L.; and Guo, B. 2022. Vector Quantized Diffusion Model for Text-to-Image Synthesis. In CVPR
2022
-
[28]
Guo, J.; Zhang, D.; Liu, X.; Zhong, Z.; Zhang, Y.; Wan, P.; and Zhang, D. 2024. LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control. arXiv preprint arXiv:2407.03168
2024 arXiv
-
[29]
Haliassos, A.; Ma, P.; Mira, R.; Petridis, S.; and Pantic, M. 2025. Jointly Learning Visual and Auditory Speech Representations from Raw Data. In ICLR
2025
-
[30]
Haliassos, A.; Mira, R.; Petridis, S.; and Pantic, M. 2022. Leveraging real talking faces via self-supervision for robust forgery detection. In CVPR
2022
-
[31]
Haliassos, A.; Zinonos, A.; Mira, R.; Petridis, S.; and Pantic, M. 2024. BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition. In ICASSP
2024
-
[32]
He, Z.; Chen, P.-Y.; and Ho, T.-Y. 2024. RIGID: A Training-free and Model-Agnostic Framework for Robust AI-Generated Image Detection. CoRR, abs/2405.20112
2024 arXiv
-
[33]
Helber, P.; Bischke, B.; Dengel, A.; and Borth, D. 2019. EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7): 2217--2226
2019
-
[34]
Hoyer, P. O. 2004. Non-negative Matrix Factorization with Sparseness Constraints. Journal of Machine Learning Research, 5: 1457--1469
2004
-
[35]
Huang, D.; and De la Torre, F. 2012. Facial Action Transfer with Personalized Bilinear Regression. In ECCV
2012
-
[36]
Huawei Noah's Ark Lab . 2022. Wukong. https://xihe.mindspore.cn/modelzoo/wukong
2022
-
[37]
Ilharco, G.; Wortsman, M.; Wightman, R.; Gordon, C.; Carlini, N.; Taori, R.; Dave, A.; Shankar, V.; Namkoong, H.; Miller, J.; Hajishirzi, H.; Farhadi, A.; and Schmidt, L. 2021. OpenCLIP
2021
-
[38]
Ji, X.; Hu, X.; Xu, Z.; Zhu, J.; Lin, C.; He, Q.; Zhang, J.; Luo, D.; Chen, Y.; Lin, Q.; Lu, Q.; and Wang, C. 2025. SONIC: Shifting Focus to Global Audio Perception in Portrait Animation. In CVPR
2025
-
[39]
J.; Wang, Q.; Shen, J.; Ren, F.; Chen, Z.; Nguyen, P.; Pang, R.; Moreno, I
Jia, Y.; Zhang, Y.; Weiss, R. J.; Wang, Q.; Shen, J.; Ren, F.; Chen, Z.; Nguyen, P.; Pang, R.; Moreno, I. L.; and Wu, Y. 2018. Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis. In NeurIPS
2018
-
[40]
Jiang, L.; Wu, W.; Li, R.-C.; Qian, C.; and Loy, C. C. 2020. DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery Detection. CVPR
2020
-
[41]
Karras, T.; Laine, S.; and Aila, T. 2019. A Style-Based Generator Architecture for Generative Adversarial Networks. In CVPR
2019
-
[42]
Khalid, H.; Tariq, S.; Kim, M.; and Woo, S. S. 2021. Fake AVC eleb: A Novel Audio-Video Multimodal Deepfake Dataset. In NeurIPS Datasets and Benchmarks Track
2021
-
[43]
A.; and Dang-Nguyen, D.-T
Khan, S. A.; and Dang-Nguyen, D.-T. 2024. CLIP ping the deception: Adapting vision-language models for universal deepfake detection. In ICMR
2024
-
[44]
Kim, T.; Choi, J.; Jeong, Y.; Noh, H.; Yoo, J.; Baek, S.; and Choi, J. 2025. Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection. In ICCV
2025
-
[45]
S.; and Noh, J
Kim, Y.; Yun, K.; Hong, S.; Cha, S.; Koo, C. S.; and Noh, J. 2026. X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection. In CVPR
2026
-
[46]
Korshunova, I.; Shi, W.; Dambre, J.; and Theis, L. 2017. Fast Face-Swap Using Convolutional Neural Networks. In ICCV
2017
-
[47]
Koutlis, C.; and Papadopoulos, S. 2024. Leveraging representations from intermediate encoder-blocks for synthetic image detection. In ECCV
2024
-
[48]
Koutlis, C.; and Papadopoulos, S. 2026. AuViRe: Audio-visual Speech Representation Reconstruction for Deepfake Temporal Localization. In WACV
2026
-
[49]
Layton, S.; De Andrade, T.; Olszewski, D.; Warren, K.; Gates, C.; Butler, K.; and Traynor, P. 2025. Every breath you don't take: Deepfake speech detection using breath. Digital Threats: Research and Practice, 6(3): 1--18
2025
-
[50]
Li, Y.; Chang, M.-C.; and Lyu, S. 2018. In ictu oculi: Exposing AI created fake videos by detecting eye blinking. In WIFS
2018
-
[51]
Li, Y.; Sun, P.; Qi, H.; and Lyu, S. 2020. Celeb-DF: A Large-scale Challenging Dataset for DeepFake Forensics . In CVPR
2020
-
[52]
Liang, Y.; Yu, M.; Li, G.; Jiang, J.; Li, B.; Yu, F.; Zhang, N.; Meng, X.; and Huang, W. 2024. Speech F orensics: Audio-Visual Speech Representation Learning for Face Forgery Detection. In NeurIPS
2024
-
[53]
Liu, H.; Tan, Z.; Tan, C.; Wei, Y.; Wang, J.; and Zhao, Y. 2024 a . Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection. In CVPR
2024
-
[54]
Liu, W.; She, T.; Liu, J.; Li, B.; Yao, D.; Liang, Z.; and Wang, R. 2024 b . Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakes. In NeurIPS
2024
-
[55]
Lopes, M. 2013. Estimating unknown sparsity in compressed sensing. In ICML
2013
-
[56]
Marra, F.; Gragnaniello, D.; Verdoliva, L.; and Poggi, G. 2019. Do GAN s leave artificial fingerprints? In Multimedia Information Processing and Retrieval
2019
-
[57]
Midjourney, Inc. 2022. Midjourney. https://www.midjourney.com/
2022
-
[58]
Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2021. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. In ICML
2021
-
[59]
Nirkin, Y.; Keller, Y.; and Hassner, T. 2019. FSGAN: Subject Agnostic Face Swapping and Reenactment. In ICCV
2019
-
[60]
Ojha, U.; Li, Y.; and Lee, Y. J. 2023. Towards Universal Fake Image Detectors that Generalize Across Generative Models. In CVPR
2023
-
[61]
Oorloff, T.; Koppisetti, S.; Bonettini, N.; Solanki, D.; Colman, B.; Yacoob, Y.; Shahriyari, A.; and Bharaj, G. 2024. AVFF : Audio-Visual Feature Fusion for Video Deepfake Detection. In CVPR
2024
-
[63]
Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. 2023. DINO v2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193
2023 arXiv
-
[64]
Park, J.; Chai, J. C. L.; Yoon, J.; and Teoh, A. B. J. 2023. Understanding the Feature Norm for Out-of-Distribution Detection. In ICCV
2023
-
[65]
Park, J.; and Owens, A. 2025. Community forensics: Using thousands of generators to train fake image detectors. In CVPR
2025
-
[66]
S.; RP, L.; Jiang, J.; et al
Perov, I.; Gao, D.; Chervoniy, N.; Liu, K.; Marangonda, S.; Um \'e , C.; Dpfks, M.; Facenheim, C. S.; RP, L.; Jiang, J.; et al. 2020. DeepFaceLab: Integrated, Flexible and Extensible Face-Swapping Framework. arXiv preprint arXiv:2005.05535
2020 arXiv
-
[67]
R.; Mukhopadhyay, R.; Namboodiri, V
Prajwal, K. R.; Mukhopadhyay, R.; Namboodiri, V. P.; and Jawahar, C. 2020. A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild. In ACM International Conference on Multimedia
2020
-
[68]
W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In ICML
2021
-
[69]
Reiss, T.; Cavia, B.; and Hoshen, Y. 2023. Detecting Deepfakes Without Seeing Any. CoRR, abs/2311.01458
2023 arXiv
-
[70]
Ricker, J.; Damm, S.; Holz, T.; and Fischer, A. 2024. Towards the detection of diffusion model deepfakes. In VISAPP
2024
-
[71]
Ricker, J.; Lukovnikov, D.; and Fischer, A. 2024. AEROBLADE : Training-Free Detection of Latent Diffusion Images Using Autoencoder Reconstruction Error. In CVPR
2024
-
[72]
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 a . High-Resolution Image Synthesis with Latent Diffusion Models. In CVPR
2022
-
[73]
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 b . High-Resolution Image Synthesis with Latent Diffusion Models. In CVPR
2022
-
[74]
R \"o ssler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; and Nie ner, M. 2019. FaceForensics++: Learning to Detect Manipulated Facial Images. ICCV
2019
-
[75]
A.; and Bhattad, A
Sarkar, A.; Mai, H.; Mahapatra, A.; Lazebnik, S.; Forsyth, D. A.; and Bhattad, A. 2024. Shadows don't lie and lines can't bend! G enerative models don't know projective geometry... for now. In CVPR
2024
-
[76]
Shi, B.; Hsu, W.; Lakhotia, K.; and Mohamed, A. 2022. Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction. In ICLR
2022
-
[77]
V.; Seitzer, M.; Baldassarre, F.; Oquab, M.; Jose, C.; Khalidov, V.; Szafraniec, M.; Yi, S.; Ramamonjisoa, M.; et al
Sim \'e oni, O.; Vo, H. V.; Seitzer, M.; Baldassarre, F.; Oquab, M.; Jose, C.; Khalidov, V.; Szafraniec, M.; Yi, S.; Ramamonjisoa, M.; et al. 2025. DINO v3. arXiv preprint arXiv:2508.10104
2025 arXiv
-
[78]
Smeu, S.; Boldisor, D.-A.; Oneata, D.; and Oneata, E. 2025. Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning localization. In CVPR
2025
-
[79]
Sun, Y.; Guo, C.; and Li, Y. 2021. ReAct : Out-of-distribution Detection With Rectified Activations. In NeurIPS
2021
-
[80]
Tan, C.; Zhao, Y.; Wei, S.; Gu, G.; Liu, P.; and Wei, Y. 2024. Frequency-aware deepfake detection: Improving generalizability through frequency space domain learning. In AAAI
2024
-
[81]
F.; and Chen, P.-Y
Tsai, C.-T.; Ko, C.-Y.; Chung, I.-H.; Wang, Y.-C. F.; and Chen, P.-Y. 2024. Understanding and Improving Training-Free AI-Generated Image Detections with Vision Foundation Models. CoRR, abs/2411.19117
2024 arXiv
-
[82]
F.; Alabdulmohsin, I.; Parthasarathy, N.; Evans, T.; Beyer, L.; Xia, Y.; Mustafa, B.; Hénaff, O.; Harmsen, J.; Steiner, A.; and Zhai, X
Tschannen, M.; Gritsenko, A.; Wang, X.; Naeem, M. F.; Alabdulmohsin, I.; Parthasarathy, N.; Evans, T.; Beyer, L.; Xia, Y.; Mustafa, B.; Hénaff, O.; Harmsen, J.; Steiner, A.; and Zhai, X. 2025. SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding...
2025 arXiv
-
[83]
T.; and Li, H
Wang, J.; Qian, X.; Zhang, M.; Tan, R. T.; and Li, H. 2023. Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert. In CVPR
2023
-
[84]
Wang, Y.; Chen, X.; Zhu, J.; Chu, W.; Tai, Y.; Wang, C.; Li, J.; Wu, Y.; Huang, F.; and Ji, R. 2021. HifiFace: 3D Shape and Semantic Prior Guided High Fidelity Face Swapping. In IJCAI
2021
-
[85]
Wei, H.; Yang, Z.; and Wang, Z. 2024. AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation. arXiv preprint arXiv:2403.17694
2024 arXiv
-
[86]
Yan, S.; Li, O.; Cai, J.; Hao, Y.; Jiang, X.; Hu, Y.; and Xie, W. 2025. A Sanity Check for AI-generated Image Detection. In ICLR
2025
-
[87]
Yang, S.; Li, H.; Wu, J.; Jing, M.; Li, L.; Ji, R.; Liang, J.; Fan, H.; and Wang, J. 2025. MegaActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer. In AAAI
2025
-
[88]
Yang, X.; Li, Y.; and Lyu, S. 2019. Exposing deep fakes using inconsistent head poses. In ICASSP
2019
-
[89]
Yu, Y.; Shin, S.; Lee, S.; Jun, C.; and Lee, K. 2023. Block Selection Method for Using Feature Norm in Out-of-Distribution Detection. In CVPR
2023
-
[90]
Zakharov, E.; Shysheya, A.; Burkov, E.; and Lempitsky, V. 2019. Few-Shot Adversarial Learning of Realistic Neural Talking Head Models. In ICCV
2019
-
[91]
Zhang, W.; Cun, X.; Wang, X.; Zhang, Y.; Shen, X.; Guo, Y.; Shan, Y.; and Wang, F. 2023. SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation. In CVPR
2023
-
[92]
Zheng, L.; Zhang, Y.; Guo, H.; Pan, J.; Tan, Z.; Lu, J.; Tang, C.; An, B.; and Yan, S. 2026. MEMO : Memory-Guided Diffusion for Expressive Talking Video Generation. TMLR
2026
-
[93]
Zhou, Y.; Han, X.; Shechtman, E.; Echevarria, J.; Kalogerakis, E.; and Li, D. 2020. MakeItTalk: Speaker-Aware Talking-Head Animation. ACM Transactions on Graphics
2020
-
[94]
Zhu, M.; Chen, H.; Yan, Q.; Huang, X.; Lin, G.; Li, W.; Tu, Z.; Hu, H.; Hu, J.; and Wang, Y. 2023. GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image. In NeurIPS
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.