REVIEW 3 major objections 5 minor 72 references
Are Foundation Models All You Need for Zero-shot Face Presentation Attack Detection?
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Frozen CLIP and DINO backbones, trained only on a one-neuron header, detect unknown face presentation attacks at rates near or above PAD models built specifically for the task, and fusing the two best backbones improves cross-database…
desk verdict Useful frozen-feature benchmark for zero-shot PAD, but the fusion gains are an in-sample selection and the SOTA claim is overstated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The framework is a frozen feature extractor plus a trainable one-neuron classification header. The extractors are the CLIP image encoder and the DINOv2 vision transformer, used as-is with pretrained weights; only the final neuron is optimized, using binary cross-entropy loss. At inference, per-frame PAD scores are averaged into a video score, and the framework optionally combines the scores of the best CLIP and DINO backbones with fixed, training-free fusion rules (MIN, MAX, SUM, AVG). The load-bearing property is that the frozen, self-supervised or contrastively trained features transfer to a top-down task they were never optimized for, so the entire adaptation cost is a single linear layer plus an off-the-shelf fusion rule.
What would settle it
A reader could rerun the cross-database and leave-one-out protocols while selecting the backbone and fusion rule on a held-out validation split, never looking at test labels; if the chosen fusion then fails to beat the stronger individual foundation model or the domain-adaptation baselines at the APCER=1% operating point, the claimed fusion improvement is an artifact of test-set selection.
Extended reading notes
Core claim
The central discovery is that the representations learned by CLIP and DINO contain enough discriminative signal to separate bona fide faces from attack presentations that the models were never trained to recognize. The paper demonstrates this by keeping the pretrained weights frozen, training only a one-neuron sigmoid header with binary cross-entropy, and evaluating under ISO/IEC 30107-3 metrics in known-attack, unknown-PAI, and cross-database protocols. On the SiW-Mv2 leave-one-out protocol, every tested foundation model outperforms the published baseline, and the best model (CLIP ViT-L-14) reaches a D-EER of 3.80% and BPCER100 of 10.54%, versus 43.34% for the baseline. In the four cross-database settings, score-level fusion (SUM or AVG) of the best-per-family backbones, DINO ViT-L-14 and CLIP ViT-L-14, reaches an average HTER of 13.15% and AUC of 93.12%, and DET-curve analysis shows lower BPCER than two state-of-the-art PAD methods at the APCER=1% operating point.
Load-bearing premise
The load-bearing premise is that selecting DINO ViT-L-14 and CLIP ViT-L-14 as the best backbones and SUM/AVG as the best fusion rule after looking at the test results would also be the best choices on unseen data, since no validation-based model selection was performed.
Editorial extensions
If this is right
- PAD subsystems can be bootstrapped from an off-the-shelf frozen backbone and a one-neuron head, removing the need to collect large labeled sets of attack presentations for each deployment domain.
- Unknown attack species, including 3D masks and makeup, become detectable without retraining: on SiW-Mv2 leave-one-out, the best foundation model reduces BPCER100 from 43.34% to 10.54% at APCER=1%.
- Combining complementary frozen representations through SUM or AVG score fusion improves cross-database performance with no development-set tuning and no additional parameters.
- Reporting only HTER and AUC can hide differences at high-security operating points; comparing APCER/BPCER reveals that the zero-shot fusion beats two domain-adaptation PAD methods at APCER=1%.
- Zero-shot foundation-model scores can be combined with traditional PAD networks, suggesting a route to improving existing systems without retraining them.
Reading between the lines
- Beyond the paper: the fusion gains are measured after the authors selected the best backbone per family and the best fusion rule by inspecting test results on each protocol, so a fairer validation-based selection would likely show smaller margins; the paper does not quantify this regression.
- Beyond the paper: because the CLIP text encoder was deliberately not used, injecting or learning text prompts that describe attack artifacts at inference is a natural next test, and could widen the gap over the image-only results.
- Beyond the paper: the same frozen-backbone-plus-linear-head recipe could be tried on iris, fingerprint, or other biometric PAD, where transferred features may also capture liveness cues without task-specific training.
- Beyond the paper: the authors note that more sophisticated fusion (boosting, bagging, weighted voting) might improve results; a testable extension is whether such learned fusions stay robust in true cross-domain deployment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies zero-shot face presentation attack detection (PAD) using frozen foundation models, specifically CLIP and DINO backbones, with only a single-neuron classification header trained on source data. It evaluates these models under known-attack, unknown-PAI, cross-database, and OULU-NPU protocols, and proposes score-level fusion between the best CLIP and DINO models. The main empirical claims are that the frozen foundation models achieve performance close to PAD-specific state-of-the-art methods, that the best single model outperforms the SiW-Mv2 leave-one-out baseline by a large margin, and that simple fusion further improves cross-database results.
Significance. If substantiated, the result is practically significant: it suggests that general-purpose frozen features can reduce the need for PAD-specific training data and can transfer across unseen attack types and databases. The paper's strengths include evaluation on five public databases, use of ISO/IEC 30107-3 metrics in addition to HTER/AUC, a simple and clearly described framework, and a public code link. However, the broadest claims about outperforming state-of-the-art PAD methods are not fully supported by the reported comparisons, and the fusion contribution depends on post-hoc model and fusion-rule selection on the test benchmarks.
major comments (3)
- [Section V-C, Tables IV and V] The fusion configuration is selected post hoc on the test benchmarks. The paper identifies DINO(ViT-L-14) and CLIP(ViT-L-14) as the best per family, and then selects SUM/AVG as the fusion rule after reporting all six backbones and four rules on the same four cross-database targets. No validation-based selection or nested protocol is described, and the best rule is not stable across targets: MAX is best on I&C&M->O while MIN beats MAX on O&C&I->M. The headline gain of 13.15% average HTER for SUM/AVG versus 16.31% for the best single model is therefore the maximum of a post-hoc search, not the expected performance of a pre-specified method. Please either add a validation-based model selection procedure or explicitly present the fusion results as exploratory and avoid claiming the selected fusion as a generalizable property of the framework.
- [Abstract and Section V-B, Table III] The abstract states that the top-performing foundation model 'outperforms by a margin the best from the state of the art observed with the leaving-one-out protocol on the SiW-Mv2 database,' but Table III compares only against the original SiW-Mv2 baseline [23]. No other state-of-the-art PAD methods are included in that protocol. The claim should be narrowed to the baseline, or additional state-of-the-art comparisons should be provided on this protocol.
- [Section V-C and Figure 4] The conclusion that the fused zero-shot framework is 'close to or even superior to' PAD-specific state-of-the-art methods is not supported by the average HTER/AUC results in Table IV. The fused method obtains 13.15% average HTER, which is worse than DADN-CDS (9.12%), CIFAS (9.57%), and FoundPAD ViT-L (9.67%). The only direct evidence of superiority is the DET-curve comparison in Figure 4, which covers two state-of-the-art methods, uses reproduced scores that the authors state may differ from the numbers in Table IV, and is read at selected operating points. Please qualify the claim accordingly or provide a comprehensive DET comparison against the full set of state-of-the-art methods.
minor comments (5)
- [Tables IV and V] The SUM and AVG rows are numerically identical. This is expected because averaging is a constant rescaling of summation, but reporting both is redundant and can confuse readers.
- [Section V-B, last paragraph] 'Most PAD techniques have poor detection performance for funny-eyes and cosmetic attacks' should read 'most foundation models' because Table III contains no comparison with other PAD techniques.
- [Section V-C and Figure 4] The text refers to 'LMDF-PAD' once; the method is LMFD-PAD elsewhere.
- [Table IV] The TransFAS and LMFD-PAD rows report identical average HTER (10.63) and AUC (94.86); please verify whether this is a typographical duplication.
- [Section III] The formatting of 'A VG' in the definition of F and in the equations should be corrected to 'AVG'.
Circularity Check
No equation-level circularity; the only circularity-adjacent issue is post-hoc selection of the fusion pair and fusion rule on the same cross-database test results.
-
other
[Section V-C (Cross-database), Table IV; also referenced in the Conclusion (Section VI).]
"we also report on the score-level fusion PAD framework ( MAX, MIN, SUM, A VG) between the best-performing model per category (i.e. DINO(ViT-L-14) and CLIP(ViT-L-14)) in Tab. IV. In particular, A VG[DINO(ViT-L-14), CLIP(ViT-L-14)] com- putes a HTER and AUC of 13.15% and 93.12%, respectively."
The backbone pair (DINO ViT-L-14 plus CLIP ViT-L-14) and the best fusion rule are selected after inspecting the same cross-database targets reported in Table IV. All six backbones and four fusion rules are evaluated on those targets, so the emphasized fusion result is the best of a post-hoc grid rather than the outcome of a pre-specified method. The reported 13.15% HTER is therefore the in-sample selection optimum, statistically forced by the selection procedure, not a derived or validated property of the framework. The per-backbone zero-shot results are unconditional and are not affected by this selection.
full rationale
The paper's core comparison is a frozen pretrained CLIP/DINO encoder plus a one-neuron header trained only on source data, with target labels unused for training; no equation in the paper defines a claimed output in terms of a fitted quantity. There is no self-definitional reduction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in by self-citation. The few self-citations (e.g., [20]-[22], [43]) are related-work references and are not load-bearing. The single-model zero-shot results on CASIA-FASD, SiW-Mv2, and the cross-database protocols stand independently of any selection. The only circularity-adjacent issue is the fusion claim: the best backbone per family and the SUM/AVG rule are chosen by comparing all variants on the same cross-database targets on which the improvement is reported, so the headline fusion gain is partly an in-sample selection artifact. Because all fusion variants are shown in Table IV, the issue is statistical rather than definitional; therefore the overall circularity score is low.
Assumptions & free parameters
free parameters (2)
- Fusion backbone selection =
DINO ViT-L-14 and CLIP ViT-L-14
- Fusion rule =
SUM/AVG
assumptions (3)
- domain assumption Pretrained weights from ImageNet/LAION-400M/LVD-142M encode features transferable to PAD without fine-tuning.
- domain assumption The classification header trained on source databases generalizes to unseen PAI species and databases.
- domain assumption The evaluation metrics (APCER, BPCER, D-EER) computed on video-level mean-fused frame scores are unbiased estimates of the true operating characteristics.
Cite this review
Pith. "Pith review of Are Foundation Models All You Need for Zero-shot Face Presentation Attack Detection?." pith.science (2026). https://pith.science/paper/QQDSDS5B
@misc{pith2026250716393,
author = {Pith},
title = {Pith review of: Are Foundation Models All You Need for Zero-shot Face Presentation Attack Detection?},
year = {2026},
howpublished = {\url{https://pith.science/paper/QQDSDS5B}},
note = {Machine review of arXiv:2507.16393}
}
read the original abstract
Although face recognition systems have undergone an impressive evolution in the last decade, these technologies are vulnerable to attack presentations (AP). These attacks are mostly easy to create and, by executing them against the system's capture device, the malicious actor can impersonate an authorised subject and thus gain access to the latter's information (e.g., financial transactions). To protect facial recognition schemes against presentation attacks, state-of-the-art deep learning presentation attack detection (PAD) approaches require a large amount of data to produce reliable detection performances and even then, they decrease their performance for unknown presentation attack instruments (PAI) or database (information not seen during training), i.e. they lack generalisability. To mitigate the above problems, this paper focuses on zero-shot PAD. To do so, we first assess the effectiveness and generalisability of foundation models in established and challenging experimental scenarios and then propose a simple but effective framework for zero-shot PAD. Experimental results show that these models are able to achieve performance in difficult scenarios with minimal effort of the more advanced PAD mechanisms, whose weights were optimised mainly with training sets that included APs and bona fide presentations. The top-performing foundation model outperforms by a margin the best from the state of the art observed with the leaving-one-out protocol on the SiW-Mv2 database, which contains challenging unknown 2D and 3D attacks
Figures
Reference graph
Works this paper leans on
-
[23]
X. Guo, Y . Liu, A. Jain, and X. Liu. Multi-domain learning for updating face anti-spoofing models. In Proc. European Conf. on Computer Vision (ECCV) , pages 230–249, 2022
work page 2022
-
[1]
S. R. Arashloo, J. Kittler, and W. Christmas. Face spoofing detection based on multiple descriptor fusion using multiscale dynamic binarized statistical image features. IEEE Trans. on Information Forensics and Security, 10(11):2396–2407, 2015
work page 2015
- [2]
-
[3]
Z. Boulkenafet, J. Komulainen, L. Li, X. Feng, and A. Hadid. Oulu- npu: A mobile face presentation attack database with real-world vari- ations. In Proc. Intl. Conf. on Automatic Face & Gesture Recognition (FG), pages 612–618, 2017
work page 2017
-
[4]
S. Chen, T. Yao, K. Zhang, Y . Chen, K. Sun, S. Ding, J. Li, F. Huang, and R. Ji. A dual-stream framework for 3d mask face presentation attack detection. In Proc. Intl. Conference on Computer Vision , pages 834–841, 2021
work page 2021
-
[5]
Z. Chen, T. Yao, K. Sheng, S. Ding, Y . Tai, J. Li, F. Huang, and X. Jin. Generalizable representation learning for mixture domain face anti-spoofing. In Proc. of the AAAI Conf. on Artificial Intelligence , volume 35, pages 1132–1139, 2021
work page 2021
-
[6]
I. Chingovska, A. Anjos, and S. Marcel. On the effectiveness of local binary patterns in face anti-spoofing. In Proc. Intl. Conf. of Biometrics Special Interest Group (BIOSIG) , pages 1–7, 2012
work page 2012
-
[7]
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei. ImageNet: A large-scale hierarchical image database. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR) , pages 248–255. Ieee, 2009
work page 2009
Show all 72 references
-
[8]
J. Deng, J. Guo, N. Xue, and S. Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR) , pages 4690–4699, 2019
2019
-
[9]
Y . S. El-Din, M. N. Moustafa, and H. Mahdi. Deep convolutional neural networks for face and iris presentation attack detection: Survey and case study. IET Biometrics, 9(5):179–193, 2020
2020
-
[10]
H. Fang, A. Liu, H. Yuan, J. Zheng, D. Zeng, Y . Liu, J. Deng, S. Escalera, X. Liu, J. Wan, et al. Unified physical-digital face attack detection. arXiv preprint arXiv:2401.17699 , 2024
2024 arXiv
-
[11]
M. Fang, H. Ali, A. Kuijper, and N. Damer. Patchswap: Boosting the generalizability of face presentation attack detection by identity-aware patch swapping. In Proc. Intl. Joint Conference on Biometrics (IJCB) , pages 1–10, 2022
2022
-
[12]
Fang and N
M. Fang and N. Damer. Face presentation attack detection by excavating causal clues and adapting embedding statistics. In Proc. Winter Conf. on Applications of Computer Vision (WCACV) , pages 6269–6279, 2024
2024
-
[13]
M. Fang, N. Damer, F. Kirchbuchner, and A. Kuijper. Learnable multi- level frequency decomposition and hierarchical attention mechanism for generalized face presentation attack detection. In Proc. Winter Conf. on Applications of Computer Vision (WCACV) , pages 3722– 3731, 2022
2022
-
[14]
M. Fang, M. Huber, and N. Damer. Synthaspoof: Developing face presentation attack detection based on privacy-friendly synthetic data. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), pages 1061–1070, 2023
2023
-
[15]
M. Fang, M. Huber, J. Fierrez, R. Ramachandra, N. Damer, A. Alkhad- dour, M. Kasantcev, V . Pryadchenko, Z. Yang, H. Huangfu, et al. Synfacepad 2023: Competition on face presentation attack detection based on privacy-aware synthetic training data. In Proc. Intl. Joint Conf. on...
2023
-
[16]
Galbally, S
J. Galbally, S. Marcel, and J. Fierrez. Biometric antispoofing methods: A survey in face recognition. IEEE Access, 2:1530–1552, 2014
2014
-
[17]
George, C
A. George, C. Ecabert, H. Shahreza, K. Kotwal, and S. Marcel. Edgeface: Efficient face recognition model for edge devices. Trans. on Biometrics, Behavior, and Identity Science (TBIOM) , 2024
2024
-
[18]
George and S
A. George and S. Marcel. Deep pixel-wise binary supervision for face presentation attack detection. In Proc. Intl. Conf. on Biometrics (ICB), pages 1–8. IEEE, 2019
2019
-
[19]
George and S
A. George and S. Marcel. On the effectiveness of vision transformers for zero-shot face anti-spoofing. In Proc. Intl. Joint Conf. on Biomet- rics (IJCB), pages 1–8, 2021
2021
-
[20]
Gonzalez-Soler, M
L. Gonzalez-Soler, M. Gomez-Barrero, and C. Busch. Toward gener- alizable facial presentation attack detection based on the analysis of facial regions. IEEE Access, 11:68512–68524, 2023
2023
-
[21]
L. J. Gonz ´alez-Soler, M. Gomez-Barrero, and C. Busch. Fisher vector encoding of dense-bsif features for unknown face presentation attack detection. In Proc. Intl. Conf. of the Biometrics Special Interest Group (BIOSIG), pages 1–6. IEEE, 2020
2020
-
[22]
L. J. Gonzalez-Soler, M. Gomez-Barrero, and C. Busch. On the generalisation capabilities of fisher vector based face presentation attack detection. IET Biometrics, 10(5):480–496, September 2021
2021
-
[24]
K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016
2016
-
[25]
Huang, Z
G. Huang, Z. Liu, L. V . D. Maaten, and K. Weinberger. Densely connected convolutional networks. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR) , pages 4700–4708, 2017
2017
-
[26]
Huang, D
H. Huang, D. Sun, Y . Liu, W. Chu, T. Xiao, J. Yuan, H. Adam, and M. Yang. Adaptive transformers for robust few-shot cross-domain face anti-spoofing. In Proc. European Conf. on Computer Vision (ECCV) , pages 37–54, 2022
2022
-
[27]
ISO/IEC 30107-3
ISO/IEC JTC1 SC37 Biometrics. ISO/IEC 30107-3. Information Technology - Biometric presentation attack detection - Part 3: Testing and Reporting. International Organization for Standardization, 2023
2023
-
[28]
C. Jia, Y . Yang, Y . Xia, Y . Chen, Z. Parekh, H. Pham, Q. Le, Y . Sung, Z. Li, and T. Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Proc. Intl. Conf. on Machine Learning (ICML) , pages 4904–4916, 2021
2021
-
[29]
Y . Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. In Proc. Intl. Conf. on Multimedia , pages 675–678, 2014
2014
-
[30]
Y . Jia, J. Zhang, S. Shan, and X. Chen. Single-side domain general- ization for face anti-spoofing. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR) , pages 8484–8493, 2020
2020
-
[31]
M. Kim, A. K. Jain, and X. Liu. Adaface: Quality adaptive margin for face recognition. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), pages 18750–18759, 2022
2022
-
[32]
Kirillov, E
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. Berg, W. Lo, et al. Segment anything. In Proc. Intl. Conf. on Computer Vision (ICCV), pages 4015–4026, 2023
2023
-
[33]
B. Koonce. Mobilenetv3. Convolutional Neural Networks with Swift for Tensorflow: Image Recognition and Dataset Categorization , pages 125–144, 2021
2021
-
[34]
Z. Li, R. Cai, H. Li, K. Lam, Y . Hu, and A. Kot. One-class knowledge distillation for face presentation attack detection. IEEE Trans. on Information Forensics and Security (TIFS) , 2022
2022
-
[35]
A. Liu, C. Zhao, Z. Yu, J. Wan, A. Su, X. Liu, Z. Tan, S. Escalera, J. Xing, Y . Liang, et al. Contrastive context-aware learning for 3d high-fidelity mask face presentation attack detection. IEEE Trans. on Information Forensics and Security (TIFS) , 2022
2022
-
[36]
Y . Liu, Y . Chen, W. Dai, C. Li, J. Zou, and H. Xiong. Causal intervention for generalizable face anti-spoofing. In Proc. Intl. Conf. on Multimedia and Expo (ICME) , pages 01–06, 2022
2022
-
[37]
Y . Liu, J. Stehouwer, A. Jourabloo, and X. Liu. Deep tree learning for zero-shot face anti-spoofing. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR) , pages 4680–4689, 2019
2019
-
[38]
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proc. Intl. Conf. on Computer Vision (ICCV) , pages 10012–10022, 2021
2021
-
[39]
Q. Meng, S. Zhao, Z. Huang, and F. Zhou. Magface: A universal representation for face recognition and quality assessment. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), pages 14225–14234, 2021
2021
-
[40]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khali- dov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. DINOv2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023
2023 arXiv
-
[41]
Ozgur, E
G. Ozgur, E. Caldeira, T. Chettaoui, F. Boutros, R. Ramachandra, and N. Damer. FoundPAD: Foundation models reloaded for face presentation attack detection. arXiv preprint arXiv:2501.02892, 2025
2025 arXiv
-
[42]
Parkhi, A
O. Parkhi, A. Vedaldi, and A. Zisserman. Deep face recognition. In Proc. British Machine Vision Conf. (BMVC) . British Machine Vision Association, 2015
2015
-
[43]
Pasmino, C
D. Pasmino, C. Aravena, J. Tapia, and C. Busch. Flickr-PAD: New face high-resolution presentation attack detection database. In Proc. Intl. Workshop on Biometrics and Forensics (IWBF), pages 1–6, 2023
2023
-
[44]
Paszke, S
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. PyTorch: An Imperative Style, High-Per...
2019
-
[45]
T. D. F. Pereira. Learning how to recognize faces in heterogeneous environments. Technical report, EPFL, 2019
2019
-
[46]
Pourpanah, M
F. Pourpanah, M. Abdar, Y . Luo, X. Zhou, R. Wang, C. Lim, X. Wang, and Q. Wu. A review of generalized zero-shot learning methods. IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI) , 45(4):4051–4070, 2022
2022
-
[47]
Radford, J
A. Radford, J. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In Proc. Intl. Conf. on Machine Learning (ICML) , pages 8748–8763, 2021
2021
-
[48]
Raghavendra and C
R. Raghavendra and C. Busch. Presentation attack detection methods for face recognition systems: A comprehensive survey. ACM Comput. Surv., 50(1):1–37, 2017
2017
-
[49]
Raghavendra, S
R. Raghavendra, S. Venkatesh, K. Raja, P. Wasnik, M. Stokkenes, and C. Busch. Fusion of multi-scale local phase quantization features for face presentation attack detection. In Proc. Intl. Conf. on Information Fusion (FUSION), pages 2107–2112, 2018
2018
-
[50]
N. Ravi, V . Gabeur, Y . Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R¨adle, C. Rolland, L. Gustafson, et al. SAM 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714 , 2024
2024 arXiv
-
[51]
Ross and N
A. Ross and N. Nandakumar. Fusion, Score-Level , pages 611–616. 2009
2009
-
[52]
Sanghvi, S
N. Sanghvi, S. Singh, A. Agarwal, M. Vatsa, and R. Singh. Mixnet for generalized face presentation attack detection. In Proc. Intl. Conf. on Pattern Recognition (ICPR) , pages 5511–5518, 2021
2021
-
[53]
Schuhmann, R
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki. LAION-400M: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021
2021 arXiv
-
[54]
Shaheed, P
K. Shaheed, P. Szczuko, M. Kumar, I. Qureshi, Q. Abbas, and I. Ullah. Deep learning techniques for biometric security: A systematic review of presentation attack detection systems. Engineering Applications of Artificial Intelligence, 129:107569, 2024
2024
-
[55]
R. Shao, X. Lan, J. Li, and P. Yuen. Multi-adversarial discriminative deep domain generalization for face presentation attack detection. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), pages 10023–10031, 2019
2019
-
[56]
R. Shao, X. Lan, and P. Yuen. Regularized fine-grained meta face anti-spoofing. In Proc. of the AAAI Conf. on Artificial Intelligence , volume 34, pages 11974–11981, 2020
2020
-
[57]
Shimizu, K
R. Shimizu, K. Asako, H. Ojima, S. Morinaga, M. Hamada, and T. Kuroda. Balanced mini-batch training for imbalanced image data classification with neural network. In Proc. Intl. Conf. on Artificial Intelligence for Industries (AI4I) , pages 27–30, 2018
2018
-
[58]
Tan and Q
M. Tan and Q. Le. EfficientNetV2: Smaller models and faster training. In Proc. Intl. Conf. on Machine Learning (ICML), pages 10096–10106, 2021
2021
-
[59]
C. Wang, Y . Lu, Y . Yang, and S. Lai. PatchNet: A simple face anti- spoofing framework via fine-grained patch recognition. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR) , pages 20281–20290, 2022
2022
-
[60]
G. Wang, H. Han, S. Shan, and X. Chen. Unsupervised adversarial domain adaptation for cross-domain face presentation attack detection. IEEE Trans. on Information Forensics and Security , 16:56–69, 2020
2020
-
[61]
K. Wang, G. Zhang, H. Yue, A. Liu, G. Zhang, H. Feng, J. Han, E. Ding, and J. Wang. Multi-domain incremental learning for face presentation attack detection. In Proc. of the AAAI Conf. on Artificial Intelligence, volume 38, pages 5499–5507, 2024
2024
-
[62]
Z. Wang, Q. Wang, W. Deng, and G. Guo. Face anti-spoofing using transformers with relation-aware mechanism. IEEE Trans. on Biometrics, Behavior, and Identity Science (TBIOM) , 4(3):439–450, 2022
2022
-
[63]
D. Wen, H. Han, and A. K. Jain. Face spoof detection with image distortion analysis. IEEE Trans. on Information Forensics and Security (TIFS), 10(4):746–761, 2015
2015
-
[64]
Z. Wu, Y . Xiong, S. Yu, and D. Lin. Unsupervised feature learning via non-parametric instance discrimination. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR) , pages 3733–3742, 2018
2018
-
[65]
Z. Xu, S. Li, and W. Deng. Learning temporal features using LSTM- CNN architecture for face anti-spoofing. In Proc. Asian Conf. on Pattern Recognition (ACPR), pages 141–145, 2015
2015
-
[66]
W. Yan, Y . Zeng, and H. Hu. Domain adversarial disentangle- ment network with cross-domain synthesis for generalized face anti- spoofing. IEEE Trans. on Circuits and Systems for Video Technology , 32(10):7033–7046, 2022
2022
-
[67]
J. Yang, Z. Lei, and S. Li. Learn convolutional neural network for face anti-spoofing. arXiv preprint arXiv:1408.5601 , 2014
2014 arXiv
-
[68]
J. Yang, Z. Lei, D. Yi, and S. Z. Li. Person-specific face antispoofing with subject domain adaptation. IEEE Trans. on Information Forensics and Security, 10(4):797–809, 2015
2015
-
[69]
Z. Yu, J. Wan, Y . Qin, X. Li, S. Li, and G. Zhao. NAS-FAS: Static- dynamic central difference network search for face anti-spoofing. IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI) , 43(9):3005–3023, 2020
2020
-
[70]
Z. Yu, C. Zhao, Z. Wang, Y . Qin, Z. Su, X. Li, F. Zhou, and Z. Zhao. Searching central difference convolutional networks for face anti-spoofing. In Proc. Intl. Conf. on Computer Vision and Pattern Recognition (CVPR), pages 5295–5305, 2020
2020
-
[71]
Zhang, Z
K. Zhang, Z. Zhang, Z. Li, and Y . Qiao. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE signal processing letters , 23(10):1499–1503, 2016
2016
-
[72]
Zhang, J
Z. Zhang, J. Yan, S. Liu, Z. Lei, D. Yi, and S. Li. A face antispoofing database with diverse attacks. In Proc. Intl. Conf. on Biometrics (ICB), pages 26–31, 2012
2012
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.