Pith. sign in

REVIEW 5 major objections 6 minor 2 cited by

Securing Social Media Against Deepfakes using Identity, Behavioral, and Geometric Signatures

T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Combining deep identity embeddings, blendshape behavior features, and golden-ratio face geometry into one 600-dimensional signature, a triplet-trained classifier, DBaGNet, is claimed to detect deepfakes across unseen manipulation types…

desk verdict The fusion idea and ablations are credible, but the headline cross-dataset gains come from an undefined 'Ours + Aug' row, so the paper needs major reporting fixes before its central claim can be trusted. read the letter →

arxiv 2412.05487 v1 pith:SOIRGMMP submitted 2024-12-07 cs.CV cs.MM

classification cs.CVcs.MM
keywords deepfakedetectionmultimediaforensicsbehavioralbiometricsfaceforgeryDBaGdescriptortripletlosscross-datasetgeneralizationidentityfeatures
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that deepfake detectors can generalize to forgeries they were never trained on if they describe a face three ways at once: a deep identity embedding of who the face is, blendshape features of how it moves and expresses, and golden-ratio-based distances that measure its proportions. The authors argue that single-cue detectors—whether they watch pixels, blinking, head pose, or heart rate—learn artifacts of specific generation algorithms and fail on unseen manipulations. They build a 600-dimensional fused descriptor, DBaG, and train a triplet-loss classifier, DBaGNet, so that real and fake videos occupy well-separated regions of an embedding space. Cross-dataset experiments, training on FaceForensics++ and testing on CelebDF, DFD, and DFDC, yield reported AUCs near 90–92%, and cross-manipulation mean AUCs exceed several state-of-the-art methods by roughly 7–16 points. If these results hold, a platform could train on known forgeries and still catch new deepfake techniques without retraining.

What carries the argument

The central object is the DBaG descriptor, a 600-dimensional per-frame signature built from three complementary cues. Identity is captured by a 512-dimensional deep embedding from a face-recognition model; behavior by 52-dimensional blendshape coefficients computed from 478 tracked facial landmarks, a representation borrowed from AR/VR avatar animation; and geometry by a 36-dimensional vector of distances from the central points of the upper and lower face to landmarks in the opposite half, a structure the authors connect to the facial golden ratio. The supporting mechanism is triplet margin loss applied to the output of a residual network with squeeze-and-excitation blocks, which places real and fake samples at separated distances in an embedding space; at test time, a new embedding's label is decided by majority vote among the closest reference embeddings.

What would settle it

Retrain the FF++-tuned model on the same split but replace the behavioral and geometric components of the DBaG vector with random values drawn once per identity and held fixed; if the cross-dataset AUC on CelebDF, DFD, and DFDC does not fall materially below the reported 89.72, 91.62, and 92.05, then the claimed generalization is not carried by the non-identity cues.

Watch

Extended reading notes

Core claim

The paper claims that a holistic facial signature combining deep identity, behavioral blendshape, and golden-ratio geometric features makes deepfake detection generalize across both unseen manipulation techniques and unseen datasets. The DBaG descriptor fuses a 512-dimensional identity embedding from a face-recognition model, a 52-dimensional behavior vector derived from facial blendshape coefficients, and a 36-dimensional vector of distances between upper- and lower-face landmark groups inspired by the facial golden ratio; 120-frame sliding windows of these vectors form the classifier input. DBaGNet, a residual network with squeeze-and-excitation blocks, maps the fused vectors into an embedding space trained with triplet margin loss, and test videos are labeled by majority vote over the nearest reference embeddings. The paper reports in-dataset AUCs above 96% on DFDC, WLDR, CelebDF, and FF++, and in the headline generalization experiment, training on FaceForensics++ and testing on CelebDF, DFD, and DFDC with augmentation yields AUCs of 89.72, 91.62, and 92.05, respectively, outperforming the compared state-of-the-art methods.

Load-bearing premise

The framework treats the face detector, landmark/blendshape extractor, and identity-embedding model as fixed, manipulation-agnostic measurement devices; if these extractors encode dataset-specific compression, generation, or identity artifacts, the cross-dataset gains would reflect artifact memorization rather than a holistic face signature.

Editorial extensions

If this is right

  • A detector trained on one known forgery family can be deployed to catch newer, unseen manipulation types without retraining, because the signature targets face properties rather than generation-specific artifacts.
  • The fusion of handcrafted cues (behavior, geometry) with a deep identity embedding gives some interpretability about which cue drove a decision, while retaining the generalization of learned representations.
  • Because classification is distance-based with a reference set, adding newly discovered forgery types to the reference set can extend detection coverage without re-running the full training pipeline.
  • The reported cross-manipulation gains of roughly 7–16 mean AUC over compared methods suggest the same three-cue signature transfers across DeepFakes, FaceSwap, and FaceShifter manipulations inside FF++.
  • Expression-swap forgeries, including the NVFAIR reenactment sets, are detected at 96–98% AUC in-dataset, indicating the descriptor covers more than face swapping.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit: the three feature extractors could be probed for dataset-specific artifacts by training the classifier on each feature type alone and comparing cross-dataset drops; the paper's ablation (identity-only → +behavior → +geometry) shows each addition helps, but not whether the extractors act as neutral sensors.
  • The nearest-neighbor voting scheme means the system's confidence depends on the balance of the reference set; a platform deploying this would need to keep the reference set balanced and current, otherwise an imbalanced real/fake reference could bias the majority vote.
  • If the golden-ratio geometric distances are the mechanism that exposes face swaps that preserve outer geometry, then those 36 distances should differ measurably between paired real and swapped videos with identical outer contours; that is directly testable.
  • The 120-frame window with 60-frame overlap sets a lower bound on detection latency on the order of seconds of video, so real-time flagging would require a streaming or shorter-window variant, which the paper does not address.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper proposes DBaGNet, a deepfake video detector that fuses three feature streams: 512-dimensional identity embeddings from AdaFace, 52-dimensional MediaPipe blendshape features, and 36-dimensional golden-ratio landmark-distance features. The 600-dimensional descriptors are sliced into 120-frame windows with 60-frame overlap and fed to a residual network with squeeze-and-excitation layers trained with triplet margin loss; at test time labels are assigned by kNN voting against reference embeddings. Experiments cover in-dataset evaluation on FF++, Celeb-DF, DFDC, DFD, WLDR, and NVFAIR, cross-manipulation evaluation within FF++, and cross-dataset evaluation from FF++ to Celeb-DF, DFD, and DFDC. The paper claims superior generalization over state-of-the-art methods.

Significance. If the cross-dataset results were reproducible, the proposed holistic descriptor and triplet-metric classifier would be a useful contribution to generalizable deepfake detection, particularly because the ablation study (Table VI) shows consistent gains from adding behavioral and geometric features to deep identity features across multiple datasets. The use of standard benchmarks and the inclusion of handcrafted, interpretable features are also positive aspects. However, the central superiority claim is currently supported only by an undocumented 'Ours + Aug' row, and one reported EER value is internally inconsistent; until these issues are resolved, the significance of the empirical claims cannot be assessed.

major comments (5)
  1. [Section IV-F, Table V] The 'Ours + Aug' row, which is the only row supporting the headline cross-dataset superiority claim, is never defined. The augmentation procedure, its hyperparameters, and whether target-dataset samples are involved are absent from Section IV-B and from every other section of the manuscript. Because the unaugmented 'Ours' row underperforms IID on Celeb-DF (82.54 vs 83.80) and on DFD (82.33 vs 93.92), all of the claimed superiority rests on an unspecified protocol, making the claim untestable and the results unreproducible.
  2. [Table V, Celeb-DF row] The 'Ours' row reports AUC 82.54 with EER 52.24 for Celeb-DF. An EER above 50 combined with an AUC well above 50 is internally inconsistent for a properly oriented score distribution, and it suggests a threshold, label, or score-direction error. This undermines confidence in the other reported metrics and should be corrected and re-verified.
  3. [Section IV-F, text vs Table V] The statement that 'the proposed model achieves superior performance over the SOTA' is contradicted by the table itself. Even the 'Ours + Aug' row is below UIA-ViT (94.68) and IID (93.92) on DFD, and the unaugmented 'Ours' row is below IID and several other baselines on two of the three datasets. The claim must be restricted to the settings actually supported by the data, or the experiments must be revised.
  4. [Tables V and VII] The 'Ours + Aug' row in Table V exactly matches the 'DIF+B+G' row in Table VII (92.05, 91.62, 89.72), yet Table VII is presented as an ablation of feature combinations with no mention of augmentation. This raises a direct reproducibility question: is the final feature combination identical to the augmented model, or does the ablation table also include the undefined augmentation? The relationship between these rows must be clarified.
  5. [Section IV-B and all result tables] No error bars, confidence intervals, multiple-seed results, or significance tests are reported. Given that some claimed differences are small (e.g., 82.54 vs 83.80 on Celeb-DF in Table V), the absence of variance information makes it impossible to determine whether the reported gains are statistically meaningful.
minor comments (6)
  1. [Equation (8)] The use of '+' to combine Fb, Fg, and Fi into a 600-dimensional vector should be replaced with an explicit concatenation operation and consistent notation.
  2. [Section III-B1] There is a typo: 'BDaG' appears where 'DBaG' is intended.
  3. [Equation (6)] The notation in Equation (6) is confusing: theta is defined as a loss expression, but the text refers to theta as an angle. The roles of the margin, scale, and angle variables should be stated clearly.
  4. [Section IV-B] The unusual train/test split for FF++ and CelebDF (first 20% of each video for training, last 20% for testing, with the remainder ignored) should be justified, since it may create temporal correlations or selection biases that affect the reported numbers.
  5. [Abstract and Section I] The sentence beginning 'Specifically, the DBaGNet classifier utilizes the extracted DBaG signatures...' is repeated verbatim in the abstract, and a similar duplication appears in Section I.
  6. [Tables III and V] The comparison tables report different metric sets; for consistency, EER should be reported for all compared methods or the omission should be explained.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the DBaG features are extracted with fixed pretrained models, the classifier is trained on ground-truth labels, and the reported gains are empirical evaluations rather than quantities forced by construction.

full rationale

The paper's central claim is an empirical generalization result. The DBaG descriptor is assembled from three externally pretrained extractors (MTCNN for face cropping, MediaPipe for landmarks and blendshapes, and AdaFace for identity embeddings), and the weights of those extractors are not fitted to the deepfake labels used in this paper. The DBaGNet classifier is trained with triplet loss on real/fake ground truth, and the cross-dataset and cross-manipulation numbers in Tables IV, V, VI, and VII are held-out evaluations, not quantities derived from the fitted parameters by construction. No equation defines a fitted constant as a prediction, and no load-bearing claim is justified by a self-citation chain. The only overlapping-author citation, HolisticDFD ([25], by Raza, Malik, and Haq), appears in the literature review and is not used to justify the proposed method or its results. The under-specified 'Ours + Aug' row in Table V is a serious reproducibility concern, but it is not circular reasoning: 'Ours + Aug' and 'Ours' are reported empirical outcomes whose experimental protocol is merely omitted. The paper therefore contains no identifiable circular step, and the circularity score is 0.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The method's load-bearing assumptions are the transferability of pretrained extractors, the informativeness of golden-ratio geometry, and identity-disjoint, protocol-matched evaluation. All are either unverified or unstated, while the main cross-dataset result depends on an undescribed augmentation procedure.

free parameters (4)
  • Triplet margin m in Eq. 9 = not reported
    Controls separation between real and fake embeddings; no value, range, or tuning schedule is given.
  • kNN neighbor count in Eq. 13 = not reported
    The final label is chosen by majority vote over an unspecified number of nearest reference embeddings; this directly affects test accuracy.
  • Augmentation procedure for 'Ours + Aug' = not described
    Headline cross-dataset AUCs in Table V come from this row, but the paper never specifies what augmentation was applied.
  • Temporal slicing window and overlap = 120 frames, 60-frame overlap
    Section III-B states this hand-chosen choice; it determines how much temporal behavioral context the model sees.
assumptions (4)
  • domain assumption Pretrained MTCNN, MobileNetV2/MLP-Mixer (MediaPipe), and AdaFace extractors provide manipulation-independent measurements without fine-tuning.
    Used for all feature extraction in Section III-B; if these models encode dataset- or artifact-specific noise, cross-dataset transfer would not be meaningful.
  • ad hoc to paper Golden-ratio-based landmark distances on lower and upper face regions carry discriminative signal about manipulated facial structure.
    Equations 3 through 5 define the geometry features, but no theory links golden-ratio proportions to deepfake artifacts; it is an aesthetic heuristic applied to forensics.
  • domain assumption The 80/20 random split for DFDC, DFD, WLDR, and NVFAIR is identity-disjoint, and the temporal 20/20 split for FF++ and CelebDF produces independent test frames.
    Section IV-B describes the splits but does not verify identity separation or frame independence; evaluation and SOTA comparisons depend on this assumption.
  • standard math Standard deep learning assumptions hold: triplet loss converges, kNN classification is consistent, and batch normalization behaves as expected.
    Implicit in Sections III-C and III-D; these are standard ML properties assumed without proof.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Securing Social Media Against Deepfakes using Identity, Behavioral, and Geometric Signatures." pith.science (2026). https://pith.science/paper/SOIRGMMP

@misc{pith2026241205487,
  author       = {Pith},
  title        = {Pith review of: Securing Social Media Against Deepfakes using Identity, Behavioral, and Geometric Signatures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SOIRGMMP}},
  note         = {Machine review of arXiv:2412.05487}
}
read the original abstract

Trust in social media is a growing concern due to its ability to influence significant societal changes. However, this space is increasingly compromised by various types of deepfake multimedia, which undermine the authenticity of shared content. Although substantial efforts have been made to address the challenge of deepfake content, existing detection techniques face a major limitation in generalization: they tend to perform well only on specific types of deepfakes they were trained on.This dependency on recognizing specific deepfake artifacts makes current methods vulnerable when applied to unseen or varied deepfakes, thereby compromising their performance in real-world applications such as social media platforms. To address the generalizability of deepfake detection, there is a need for a holistic approach that can capture a broader range of facial attributes and manipulations beyond isolated artifacts. To address this, we propose a novel deepfake detection framework featuring an effective feature descriptor that integrates Deep identity, Behavioral, and Geometric (DBaG) signatures, along with a classifier named DBaGNet. Specifically, the DBaGNet classifier utilizes the extracted DBaG signatures, leveraging a triplet loss objective to enhance generalized representation learning for improved classification. Specifically, the DBaGNet classifier utilizes the extracted DBaG signatures and applies a triplet loss objective to enhance generalized representation learning for improved classification. To test the effectiveness and generalizability of our proposed approach, we conduct extensive experiments using six benchmark deepfake datasets: WLDR, CelebDF, DFDC, FaceForensics++, DFD, and NVFAIR. Specifically, to ensure the effectiveness of our approach, we perform cross-dataset evaluations, and the results demonstrate significant performance gains over several state-of-the-art methods.

Figures

Figures reproduced from arXiv: 2412.05487 by the authors.

Figure 1
Figure 1. Analysis of in- and cross-dataset representations of individual [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Visual representation of the proposed feature descriptor, combining (a) deep identity features, (b) behavioral features, and (c) face geometry features [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Detailed overview of the proposed feature descriptor. In preprocessing step, face detection and cropping are performed followed by feature extraction [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Detailed architecture of the proposed DBaGNet with triplet loss for representation learning. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Effectiveness of proposed framework on cross-manipulation evaluation: (a) shows the effectiveness of model when trained on DeepFake (DF) and [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SHIELD: A Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks

    cs.SD 2025-07 conditional novelty 5.0 of 10

    The paper claims that a collaborative defense generator plus triplet learning keeps audio deepfake detectors at roughly 98 percent accuracy against GAN-based anti-forensic attacks.

  2. Transferable Adversarial Attacks on Audio Deepfake Detection

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A transferable GAN-based attack that preserves transcription and perceptual quality can substantially degrade current audio deepfake detection systems.

Reference graph

Works this paper leans on

67 extracted references · 41 canonical work pages · cited by 2 Pith papers

  1. [1]

    Behavior modeling and forensics for multimedia social networks,

    H. V . Zhao, W. S. Lin, and K. R. Liu, “Behavior modeling and forensics for multimedia social networks,” IEEE Signal Processing Magazine , vol. 26, no. 1, pp. 118–139, 2009

  2. [2]

    Automatic face reenactment,

    P. Garrido, L. Valgaerts, O. Rehmsen, T. Thormaehlen, P. Perez, and C. Theobalt, “Automatic face reenactment,” in 2014 IEEE Conference on Computer Vision and Pattern Recognition . IEEE, June 2014. [Online]. Available: http://dx.doi.org/10.1109/CVPR.2014.537 JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, DECEMBER 2024 11

  3. [3]

    Real-time expression transfer for facial reenactment,

    J. Thies, M. Zollh ¨ofer, M. Nießner, L. Valgaerts, M. Stamminger, and C. Theobalt, “Real-time expression transfer for facial reenactment,” ACM Trans. Graph. , vol. 34, no. 6, nov 2015. [Online]. Available: https://doi.org/10.1145/2816795.2818056

  4. [4]

    Securing the socio-cyber world: multiorder attribute node association classification for manip- ulated media,

    S. Xiao, G. Lan, J. Yang, Y . Li, and J. Wen, “Securing the socio-cyber world: multiorder attribute node association classification for manip- ulated media,” IEEE Transactions on Computational Social Systems , no. 99, pp. 1–10, 2022

  5. [5]

    Automated face swapping and its detection,

    Y . Zhang, L. Zheng, and V . L. Thing, “Automated face swapping and its detection,” in 2017 IEEE 2nd international conference on signal and image processing (ICSIP) . IEEE, 2017, pp. 15–19

  6. [6]

    Deepfake source detection in a heart beat,

    U. A. C ¸ iftc ¸i,˙I. Demir, and L. Yin, “Deepfake source detection in a heart beat,” The Visual Computer , vol. 40, no. 4, pp. 2733–2750, 2024

  7. [7]

    Deepvision: Deepfakes detection using human eye blinking pattern,

    T. Jung, S. Kim, and K. Kim, “Deepvision: Deepfakes detection using human eye blinking pattern,” IEEE Access , vol. 8, pp. 83 144–83 154, 2020

  8. [8]

    In ictu oculi: Exposing ai gen- erated fake face videos by detecting eye blinking,

    Y . Li, M.-C. Chang, and S. Lyu, “In ictu oculi: Exposing ai gen- erated fake face videos by detecting eye blinking,” arXiv preprint arXiv:1806.02877, 2018

Show all 67 references
  1. [10]

    Exposing deepfake videos by detecting face warping artif acts,

    Y . Li, “Exposing deepfake videos by detecting face warping artif acts,” arXiv preprint arXiv:1811.00656 , 2018

  2. [11]

    Predicting heart rate variations of deepfake videos using neural ode,

    S. Fernandes, S. Raj, E. Ortiz, I. Vintila, M. Salter, G. Urosevic, and S. Jha, “Predicting heart rate variations of deepfake videos using neural ode,” in Proceedings of the IEEE/CVF international conference on computer vision workshops , 2019, pp. 0–0

  3. [12]

    Msta-net: Forgery detection by generating manipulation trace based on multi-scale self- texture attention,

    J. Yang, S. Xiao, A. Li, W. Lu, X. Gao, and Y . Li, “Msta-net: Forgery detection by generating manipulation trace based on multi-scale self- texture attention,” IEEE transactions on circuits and systems for video technology, vol. 32, no. 7, pp. 4854–4866, 2021

  4. [13]

    Recurrent convolutional strategies for face manipulation detection in videos,

    E. Sabir, J. Cheng, A. Jaiswal, W. AbdAlmageed, I. Masi, and P. Natara- jan, “Recurrent convolutional strategies for face manipulation detection in videos,” Interfaces (GUI), vol. 3, no. 1, pp. 80–87, 2019

  5. [15]

    Exposing deep fakes using inconsistent head poses,

    X. Yang, Y . Li, and S. Lyu, “Exposing deep fakes using inconsistent head poses,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 8261–8265

  6. [16]

    Exploiting visual artifacts to expose deepfakes and face manipulations,

    F. Matern, C. Riess, and M. Stamminger, “Exploiting visual artifacts to expose deepfakes and face manipulations,” in 2019 IEEE Winter Applications of Computer Vision Workshops (WACVW) . IEEE, 2019, pp. 83–92

  7. [17]

    Protecting world leaders against deep fakes

    S. Agarwal, H. Farid, Y . Gu, M. He, K. Nagano, and H. Li, “Protecting world leaders against deep fakes.” in CVPR workshops , vol. 1, 2019, p. 38

  8. [18]

    Detecting gan-generated imagery using saturation cues,

    S. McCloskey and M. Albright, “Detecting gan-generated imagery using saturation cues,” in 2019 IEEE international conference on image processing (ICIP). IEEE, 2019, pp. 4584–4588

  9. [19]

    Deeprhythm: Exposing deepfakes with attentional visual heartbeat rhythms,

    H. Qi, Q. Guo, F. Juefei-Xu, X. Xie, L. Ma, W. Feng, Y . Liu, and J. Zhao, “Deeprhythm: Exposing deepfakes with attentional visual heartbeat rhythms,” in Proceedings of the 28th ACM international conference on multimedia, 2020, pp. 4318–4327

  10. [20]

    Towards robust interpretability with self-explaining neural networks,

    D. Alvarez-Melis and T. S. Jaakkola, “Towards robust interpretability with self-explaining neural networks,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , ser. NIPS’18. Red Hook, NY , USA: Curran Associates Inc., 2018, p. 7786–7795

  11. [21]

    The bayesian case model: a generative approach for case-based reasoning and prototype classification,

    B. Kim, C. Rudin, and J. Shah, “The bayesian case model: a generative approach for case-based reasoning and prototype classification,” in Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2 , ser. NIPS’14. Cambridge, MA, USA: MI...

  12. [22]

    C. Chen, O. Li, C. Tao, A. J. Barnett, J. Su, and C. Rudin, This looks like that: deep learning for interpretable image recognition. Red Hook, NY , USA: Curran Associates Inc., 2019

  13. [23]

    D-tsm: Discriminative temporal shift module for action recognition,

    S. Lee and S. Hong, “D-tsm: Discriminative temporal shift module for action recognition,” in 2023 20th International Conference on Ubiquitous Robots (UR) . IEEE, 2023, pp. 133–136

  14. [24]

    Modulation of visual contrast sensitivity with trns across the visual system, evidence from stimulation and simulation,

    W. Potok, A. Post, V . Beliaeva, M. B¨achinger, A. M. Cassar`a, E. Neufeld, R. Polania, D. Kiper, and N. Wenderoth, “Modulation of visual contrast sensitivity with trns across the visual system, evidence from stimulation and simulation,” eneuro, vol. 10, no. 6, 2023

  15. [25]

    Holisticdfd: Infusing spatiotemporal transformer embeddings for deepfake detection,

    M. A. Raza, K. M. Malik, and I. U. Haq, “Holisticdfd: Infusing spatiotemporal transformer embeddings for deepfake detection,” Infor- mation Sciences, vol. 645, p. 119352, 2023

  16. [26]

    Detecting deepfake video by learning two-level features with two-stream convolutional neural network,

    Z. Zhao, P. Wang, and W. Lu, “Detecting deepfake video by learning two-level features with two-stream convolutional neural network,” in Proceedings of the 2020 6th International Conference on Computing and Artificial Intelligence , 2020, pp. 291–297

  17. [27]

    Digital image forgery detection using deep autoencoder and cnn features,

    S. Bibi, A. Abbasi, I. U. Haq, S. W. Baik, and A. Ullah, “Digital image forgery detection using deep autoencoder and cnn features,” Hum. Cent. Comput. Inf. Sci , vol. 11, pp. 1–17, 2021

  18. [28]

    Multi-task learning for detecting and segmenting manipulated facial images and videos,

    H. H. Nguyen, F. Fang, J. Yamagishi, and I. Echizen, “Multi-task learning for detecting and segmenting manipulated facial images and videos,” in 2019 IEEE 10th international conference on biometrics theory, applications and systems (BTAS) . IEEE, 2019, pp. 1–8

  19. [29]

    Using cascade cnn-lstm-fcns to identify ai-altered video based on eye state sequence,

    M. S. Saealal, M. Z. Ibrahim, D. J. Mulvaney, M. I. Shapiai, and N. Fadilah, “Using cascade cnn-lstm-fcns to identify ai-altered video based on eye state sequence,” PLoS One, vol. 17, no. 12, p. e0278989, 2022

  20. [30]

    Gan is a friend or foe? a framework to detect various fake face images,

    S. Tariq, S. Lee, H. Kim, Y . Shin, and S. S. Woo, “Gan is a friend or foe? a framework to detect various fake face images,” in Proceedings of the 34th ACM/SIGAPP Symposium on Applied Computing , 2019, pp. 1296–1303

  21. [31]

    Pudd: Towards robust multi-modal prototype-based deepfake detection,

    A. L. Pellicer, Y . Li, and P. Angelov, “Pudd: Towards robust multi-modal prototype-based deepfake detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 3809–3817

  22. [32]

    Deepfake catcher: Can a simple fusion be effective and outperform complex dnns?

    A. Agarwal and N. Ratha, “Deepfake catcher: Can a simple fusion be effective and outperform complex dnns?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 3791–3801

  23. [33]

    Exposing deepfake videos by detecting face warping artifacts,

    Y . Li and S. Lyu, “Exposing deepfake videos by detecting face warping artifacts,” in CVPR Workshops , 2018. [Online]. Available: https://api.semanticscholar.org/CorpusID:53298495

  24. [34]

    Face x-ray for more general face forgery detection,

    L. Li, J. Bao, T. Zhang, H. Yang, D. Chen, F. Wen, and B. Guo, “Face x-ray for more general face forgery detection,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 5000–5009, 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:...

  25. [35]

    Spatial-phase shallow learning: Rethinking face forgery detection in frequency domain,

    H. Liu, X. Li, W. Zhou, Y . Chen, Y . He, H. Xue, W. Zhang, and N. Yu, “Spatial-phase shallow learning: Rethinking face forgery detection in frequency domain,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 772–781, 2021. [Online]. Available: ...

  26. [36]

    Lips don’t lie: A generalisable and robust approach to face forgery detection,

    A. Haliassos, K. V ougioukas, S. Petridis, and M. Pantic, “Lips don’t lie: A generalisable and robust approach to face forgery detection,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5037–5047, 2020. [Online]. Available: https://api.semantic...

  27. [37]

    Generalizing face forgery detection with high-frequency features,

    Y . Luo, Y . Zhang, J. Yan, and W. Liu, “Generalizing face forgery detection with high-frequency features,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 16 312–16 321,

  28. [38]

    Transcending forgery specificity with latent space augmentation for generalizable deepfake detection,

    Z. Yan, Y . Luo, S. Lyu, Q. Liu, and B. Wu, “Transcending forgery specificity with latent space augmentation for generalizable deepfake detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 8984–8994

  29. [39]

    Gm-df: Generalized multi-scenario deepfake detection,

    Y . Lai, Z. Yu, J. Yang, B. Li, X. Kang, and L. Shen, “Gm-df: Generalized multi-scenario deepfake detection,” arXiv preprint arXiv:2406.20078 , 2024

  30. [40]

    Joint face detection and alignment using multitask cascaded convolutional networks,

    K. Zhang, Z. Zhang, Z. Li, and Y . Qiao, “Joint face detection and alignment using multitask cascaded convolutional networks,” IEEE signal processing letters , vol. 23, no. 10, pp. 1499–1503, 2016

  31. [41]

    Automated blendshape per- sonalization for faithful face animations using commodity smartphones,

    T. Menzel, M. Botsch, and M. E. Latoschik, “Automated blendshape per- sonalization for faithful face animations using commodity smartphones,” in Proceedings of the 28th ACM Symposium on Virtual Reality Software and Technology, 2022, pp. 1–9

  32. [42]

    Mediapipe face landmarker,

    Google, “Mediapipe face landmarker,” https://storage.googleapis.com/ mediapipe-assets/Model%20Card%20MediaPipe%20Face%20Mesh% 20V2.pdf, 2024

  33. [43]

    G. B. Meisner, The golden ratio: The divine beauty of mathematics . Race Point Publishing, 2018

  34. [44]

    Adaface: Quality adaptive margin for face recognition,

    M. Kim, A. K. Jain, and X. Liu, “Adaface: Quality adaptive margin for face recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 18 750–18 759. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, DECEMBER 2024 12

  35. [45]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778

  36. [46]

    Squeeze-and-excitation networks,

    J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132–7141

  37. [47]

    The deepfake detection challenge (dfdc) dataset,

    B. Dolhansky, J. Bitton, B. Pflaum, J. Lu, R. Howes, M. Wang, and C. C. Ferrer, “The deepfake detection challenge (dfdc) dataset,” arXiv preprint arXiv:2006.07397, 2020

  38. [48]

    Celeb-df: A large-scale challenging dataset for deepfake forensics,

    Y . Li, X. Yang, P. Sun, H. Qi, and S. Lyu, “Celeb-df: A large-scale challenging dataset for deepfake forensics,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2020

  39. [49]

    Contributing data to deepfake detection research,

    N. Dufour and A. Gully, “Contributing data to deepfake detection research,” https://ai.googleblog.com/2019/09/ contributing-data-to-deepfake-detection.html, 2019

  40. [50]

    Faceforensics++: Learning to detect manipulated facial images,

    A. Rossler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nießner, “Faceforensics++: Learning to detect manipulated facial images,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 1–11

  41. [51]

    Avatar fingerprinting for authorized use of synthetic talking-head videos,

    E. Prashnani, K. Nagano, S. D. Mello, D. Luebke, and O. Gallo, “Avatar fingerprinting for authorized use of synthetic talking-head videos,” arXiv preprint arXiv:2305.03713, 2023

  42. [52]

    Detecting deep- fake videos from appearance and behavior,

    S. Agarwal, H. Farid, T. El-Gaaly, and S.-N. Lim, “Detecting deep- fake videos from appearance and behavior,” in 2020 IEEE international workshop on information forensics and security (WIFS) . IEEE, 2020, pp. 1–6

  43. [53]

    Two-stream neural networks for tampered face detection,

    P. Zhou, X. Han, V . I. Morariu, and L. S. Davis, “Two-stream neural networks for tampered face detection,” in 2017 IEEE conference on computer vision and pattern recognition workshops (CVPRW) . IEEE, 2017, pp. 1831–1839

  44. [54]

    Mesonet: a compact facial video forgery detection network,

    D. Afchar, V . Nozick, J. Yamagishi, and I. Echizen, “Mesonet: a compact facial video forgery detection network,” in 2018 IEEE international workshop on information forensics and security (WIFS) . IEEE, 2018, pp. 1–7

  45. [55]

    Multi- attentional deepfake detection,

    H. Zhao, W. Zhou, D. Chen, T. Wei, W. Zhang, and N. Yu, “Multi- attentional deepfake detection,” in Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition , 2021, pp. 2185–2194

  46. [56]

    Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection,

    L. Chen, Y . Zhang, Y . Song, L. Liu, and J. Wang, “Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 18 710–18 719

  47. [57]

    Face x-ray for more general face forgery detection,

    L. Li, J. Bao, T. Zhang, H. Yang, D. Chen, F. Wen, and B. Guo, “Face x-ray for more general face forgery detection,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 5001–5010

  48. [58]

    Learning self- consistency for deepfake detection,

    T. Zhao, X. Xu, M. Xu, H. Ding, Y . Xiong, and W. Xia, “Learning self- consistency for deepfake detection,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 15 023–15 033

  49. [59]

    Detecting deepfakes with self-blended images,

    K. Shiohara and T. Yamasaki, “Detecting deepfakes with self-blended images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 720–18 729

  50. [60]

    Efficientnet: Rethinking model scaling for con- volutional neural networks,

    M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for con- volutional neural networks,” in International conference on machine learning. PMLR, 2019, pp. 6105–6114

  51. [61]

    Generalizing face forgery detec- tion with high-frequency features,

    Y . Luo, Y . Zhang, J. Yan, and W. Liu, “Generalizing face forgery detec- tion with high-frequency features,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 16 317–16 326

  52. [62]

    Dual contrastive learning for general face forgery detection,

    K. Sun, T. Yao, S. Chen, S. Ding, J. Li, and R. Ji, “Dual contrastive learning for general face forgery detection,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 2, 2022, pp. 2316– 2324

  53. [63]

    Implicit identity driven deepfake face swapping detection,

    B. Huang, Z. Wang, J. Yang, J. Ai, Q. Zou, Q. Wang, and D. Ye, “Implicit identity driven deepfake face swapping detection,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 4490–4499

  54. [64]

    Learning to generalize: Meta-learning for domain generalization,

    D. Li, Y . Yang, Y .-Z. Song, and T. Hospedales, “Learning to generalize: Meta-learning for domain generalization,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018

  55. [65]

    F 3net: fusion, feedback and focus for salient object detection,

    J. Wei, S. Wang, and Q. Huang, “F 3net: fusion, feedback and focus for salient object detection,” in Proceedings of the AAAI conference on artificial intelligence, vol. 34, no. 07, 2020, pp. 12 321–12 328

  56. [66]

    Domain general face forgery detection by learning to weight,

    K. Sun, H. Liu, Q. Ye, Y . Gao, J. Liu, L. Shao, and R. Ji, “Domain general face forgery detection by learning to weight,” in Proceedings of the AAAI conference on artificial intelligence , vol. 35, no. 3, 2021, pp. 2638–2646

  57. [67]

    Local relation learn- ing for face forgery detection,

    S. Chen, T. Yao, Y . Chen, S. Ding, J. Li, and R. Ji, “Local relation learn- ing for face forgery detection,” in Proceedings of the AAAI conference on artificial intelligence , vol. 35, no. 2, 2021, pp. 1081–1088

  58. [68]

    Uia-vit: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection,

    W. Zhuang, Q. Chu, Z. Tan, Q. Liu, H. Yuan, C. Miao, Z. Luo, and N. Yu, “Uia-vit: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection,” in European conference on computer vision . Springer, 2022, pp. 391–407. Muhammad Umar Farooq is c...

  59. [2021]

    Available: https://api.semanticscholar.org/CorpusID: 232320599

    [Online]. Available: https://api.semanticscholar.org/CorpusID: 232320599

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.