REVIEW 3 major objections 5 minor 2 cited by
Benchmarking Foundation Models for Zero-Shot Biometric Tasks
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A benchmark of 41 vision-language models finds that frozen embeddings verify faces and irises at up to 97% without fine-tuning.
desk verdict A broad and useful zero-shot biometric benchmark that needs a leakage audit before its headline 'zero-shot' claims are fully supported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the frozen vision encoder of each foundation model acting as a generic feature extractor. The paper takes the vision-tower embeddings from models such as CLIP, OpenCLIP, BLIP-2, DINO, DINOv2, and InternVL3, then either compares pairs by cosine similarity for verification or trains a shallow head such as logistic regression, SVM, LDA, KNN, or a two-layer MLP on the frozen features for classification and attack detection. Since no model weights are updated, the benchmark isolates what the pretrained representations alone carry over to biometric tasks.
What would settle it
An audit that checks LFW, IITD iris, FaceForensics++, and VMER images against the pretraining corpora of the top models: if near-duplicates are found and removing them changes the reported TMR@1%FMR by more than a few points, then the zero-shot result is substantially memorization rather than transfer.
Extended reading notes
Core claim
The central claim is that the image encoders of pretrained VLMs and MLLMs, used without any fine-tuning, already encode discriminative biometric information. For verification, the paper extracts the frozen embeddings of each image and computes cosine similarity between pairs; this yields the strong LFW and IITD-R results. For attribute prediction and attack detection, a lightweight classifier head is trained on the frozen features, producing near-perfect gender and iris PAD accuracies and roughly 90% accuracy for deepfake detection with InternVL3-78B on FaceForensics++. The paper interprets these results as evidence that language-supervised and self-supervised pretraining can substitute for task-specific metric learning in at least some biometric pipelines, with performance depending strongly on model family and task difficulty.
Load-bearing premise
The load-bearing premise is that the benchmark images were not part of the models' pretraining data, so the reported zero-shot scores reflect transferable discrimination rather than memorized identities; the paper does not audit for this overlap.
Editorial extensions
If this is right
- Face verification can be deployed without biometric-specific training: OpenCLIP-H/14 reaches 96.77% TMR@1%FMR on LFW using only cosine similarity on frozen embeddings.
- Iris recognition transfers too: DINO-ViT-B/16 reaches 97.55% TMR@1%FMR on IITD-R-Full, with cropped and full iris images performing similarly for many models.
- Lightweight biometric pipelines are possible: frozen embeddings plus simple classifiers exceed 99% accuracy for gender classification and iris presentation attack detection with the best models.
- Deepfake detection benefits from large MLLM embeddings, with InternVL3-78B reaching nearly 90% accuracy with a simple MLP head on FaceForensics++.
- Morph attack detection is the outlier: several models stay at or near chance level, indicating that fine-grained morph artifacts are not captured by current frozen embeddings.
Reading between the lines
- My inference: the text encoders and chat interfaces of these models are largely unexploited here, so a natural extension is to prompt MLLMs directly with identity and attribute questions and compare those answers against the embedding-only scores.
- My inference: the large performance gap between IITD and UND iris datasets, and between LFW and CPLFW, points to cross-sensor and cross-pose shifts as the main remaining bottleneck, which could be tested by measuring how frozen-embedding verification degrades on unseen sensors and demographics.
- My inference: because the pretraining corpora are web-scale, a direct deduplication audit of benchmark images against training data would settle whether the top scores reflect transferable discrimination or memorization, and such an audit is easy to run.
- My inference: the fact that different model families rank similarly across tasks suggests that pretraining objective and data distribution, rather than parameter count, drive biometric transferability, which a controlled pretraining ablation could confirm.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a broad benchmark of 41 publicly available vision-language and multimodal models on six biometric tasks in the face and iris domains: face verification, gender and ethnicity classification, iris recognition, iris presentation attack detection (PAD), deepfake detection, and morph attack detection. Using frozen image encoders, the authors compute cosine-similarity verification scores or train lightweight classifiers on top of extracted embeddings. Headline results include 96.77% TMR@1%FMR on LFW with OpenCLIP-H/14, 97.55% on IITD-R-Full with DINO-ViT-B/16, near-99% gender classification with several models, and strong PAD accuracy with DINO/DINOv2 embeddings. The paper also reports negative results, such as chance-level morph detection with Chameleon.
Significance. If the zero-shot interpretation is established, this is a useful reference benchmark for an important question: whether generic foundation-model embeddings carry transferable biometric discriminability. The breadth across 41 models and six tasks is a genuine contribution, and the consistency of the main verification tables makes the raw measurements credible. The paper also gives credit where due: it uses standard public datasets and protocols, reports both strengths and failures across model families, and makes no parameter-free derivation claims that could be circular. The main value is as an empirical baseline; the main risk is that the zero-shot label is not yet supported by a leakage audit.
major comments (3)
- [V-D and VII-5] The central 'zero-shot' interpretation is not yet supported because the manuscript never audits whether benchmark images appeared in pretraining corpora. OpenCLIP-H/14 is trained on LAION-5B, DINO and DINOv2 on large web-curated image sets, and LFW, CFP, and AgeDB contain public internet photographs of identifiable individuals; near-duplicate or membership information could plausibly inflate the reported TMR@1%FMR values such as 96.77% on LFW. The paper itself notes 'implicit exposure to ocular patterns during pretraining' in Section V-D and lists understanding the training data as future work in Section VII-5, so the issue is acknowledged but not resolved. Please add a contamination analysis: retrieval-based near-duplicate checks against accessible training corpora, membership or rank statistics, or at minimum a quantitative argument for each high-performing model. Without this, the headline numbers do not distinguish generalization from memorization.
- [III-D and Table IV] The zero-shot PAD result ('Acc' column) is not reproducible because no decision rule is specified. Table IV reports binary classification accuracy between Patterned and Normal iris images under 'zero-shot inference,' but the paper never states how a class label is derived from the embeddings: threshold on cosine similarity to class prototypes, nearest-centroid assignment, text-prompt similarity, or another rule. Please specify the protocol, including threshold selection and class priors, or relabel the column so that it does not imply an unsupervised decision rule.
- [III-B1, IV-B, and Table II] All attribute-classification results are reported without variance: Section III-B1 states 'Standard deviations are omitted for brevity,' and Section IV-B collapses 5-fold cross-validation to a single mean. Many models differ by less than 0.2% in gender accuracy (values cluster near 99.9%), so the ranking claims in Section V-C are not statistically supported. Please report standard deviations, fold-level extrema, or confidence intervals for Table II.
minor comments (5)
- [V-A] The text refers to a '40-year protocol' and assigns the best 30-year result (21.23%) to LLaVA-1.5; Table I has only 5/10/20/30-year columns. Also, the claim that 'BLIP [7] outperformed all others' on CPLFW contradicts Table I, where BLIP-Base scores 7.60 and BLIP-Large 6.03 while BLIP2-t5-xxl scores 43.03.
- [V-A] The best CFP-FP result is attributed to 'CLIP-L/14 [28],' but [28] is a prompt-tuning paper and Table I shows the 87.63% value under CLIP-L-32, not a CLIP-L/14 row. Please correct the model name and citation.
- [Figure 5] The LLaVA architecture figure cites reference [8] but should cite reference [10]; the caption also contains the typo 'We use only use.'
- [Figure 10] The caption contains the typo 'socres' and should read 'scores.'
- [IV-C] The morph attack detection setup says a decision tree classifier is trained for each model, but Table VI does not report decision-tree specifics; please state the tree hyperparameters, whether results are averaged over runs, and how BPCER@10%APCER thresholds were selected.
Circularity Check
No significant circularity: the paper's claims are empirical benchmark measurements with no derivational chain that reduces to fitted values or load-bearing self-citations.
full rationale
The paper is an empirical benchmarking study: face and iris verification scores are computed by cosine similarity between frozen VLM embeddings without any training (Section IV-A, V-A, V-D), so the headline TMR values (e.g., OpenCLIP-H/14's 96.77% on LFW, DINO-ViT-B/16's 97.55% on IITD-R-Full) are direct measurements of off-the-shelf representations, not outputs of a fitted model. The gender/ethnicity, PAD, DeepFake, and morph experiments train lightweight classifiers on extracted embeddings, but they do so under disjoint identity splits and clearly distinguish zero-shot inference from classifier-based evaluation; the classifier results are reported as few-shot/classifier results, not disguised predictions of the frozen encoders. The only self-citations (Iris-SAM, ChatGPT iris analysis, and the two demorphing papers by Shukla and Ross) are used for context or to justify a morph-training strategy, and the strategy itself is fully described in the paper (Section III-F: creating 5,000 training morphs from SMDD with OpenCV/dlib), so no load-bearing claim is imported from an unverified self-citation. The reader's expressed concern about dataset leakage into pretraining corpora is a real external-validity risk for the term 'zero-shot,' but it is not a circularity step: no equation, fitted parameter, or self-citation makes the reported numbers equivalent to the paper's inputs by construction. The central claims are self-contained empirical measurements against fixed public benchmarks, so the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Publicly released model checkpoints faithfully represent the published architectures and weights used in the benchmark.
- domain assumption Standard verification protocols for LFW, CFP, CPLFW, AgeDB, and iris datasets are applied as originally defined.
- domain assumption Benchmark test images are not present in the pretraining corpora of the evaluated models.
- domain assumption Cosine similarity is a valid fixed metric for comparing embeddings across all model families.
Cite this review
Pith. "Pith review of Benchmarking Foundation Models for Zero-Shot Biometric Tasks." pith.science (2026). https://pith.science/paper/7NC6CQH4
@misc{pith2026250524214,
author = {Pith},
title = {Pith review of: Benchmarking Foundation Models for Zero-Shot Biometric Tasks},
year = {2026},
howpublished = {\url{https://pith.science/paper/7NC6CQH4}},
note = {Machine review of arXiv:2505.24214}
}
read the original abstract
The advent of foundation models, particularly Vision-Language Models (VLMs) and Multi-modal Large Language Models (MLLMs), has redefined the frontiers of artificial intelligence, enabling remarkable generalization across diverse tasks with minimal or no supervision. Yet, their potential in biometric recognition and analysis remains relatively underexplored. In this work, we introduce a comprehensive benchmark that evaluates the zero-shot and few-shot performance of state-of-the-art publicly available VLMs and MLLMs across six biometric tasks spanning the face and iris modalities: face verification, soft biometric attribute prediction (gender and race), iris recognition, presentation attack detection (PAD), and face manipulation detection (morphs and deepfakes). A total of 41 VLMs were used in this evaluation. Experiments show that embeddings from these foundation models can be used for diverse biometric tasks with varying degrees of success. For example, in the case of face verification, a True Match Rate (TMR) of 96.77 percent was obtained at a False Match Rate (FMR) of 1 percent on the Labeled Face in the Wild (LFW) dataset, without any fine-tuning. In the case of iris recognition, the TMR at 1 percent FMR on the IITD-R-Full dataset was 97.55 percent without any fine-tuning. Further, we show that applying a simple classifier head to these embeddings can help perform DeepFake detection for faces, Presentation Attack Detection (PAD) for irides, and extract soft biometric attributes like gender and ethnicity from faces with reasonably high accuracy. This work reiterates the potential of pretrained models in achieving the long-term vision of Artificial General Intelligence.
Figures
Figures from the paper (11 more)
Forward citations
Cited by 2 Pith papers
-
DiscoGen: Procedural Generation of Algorithm Discovery Tasks in Machine Learning
DiscoGen procedurally generates billions of configurable ML algorithm-discovery tasks and a fixed DiscoBench subset so algorithm-discovery agents can be trained and evaluated without contamination or saturation.
-
Foundation versus Domain-specific Models: Performance Comparison, Fusion, and Explainability in Face Recognition
On standard face benchmarks, domain-specific face recognition models beat zero-shot foundation models, adding context or fusing scores improves performance at low false-match rates, and GPT-4o can explain and sometime...
Reference graph
Works this paper leans on
-
[1]
On the Opportunities and Risks of Foundation Models,
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill et al. , “On the Opportunities and Risks of Foundation Models,” arXiv preprint arXiv:2108.07258, 2021
arXiv 2021
-
[2]
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “GPT-4 Technical Report,” arXiv preprint arXiv:2303.08774 , 2023
arXiv 2023
-
[3]
Segment anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollar, and R. Girshick, “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023, pp. 4015–4026
2023
-
[4]
Segment anything in medical images,
J. Ma, Y . He, F. Li, L. Han, C. You, and B. Wang, “Segment anything in medical images,” Nature Communications, vol. 15, no. 1, p. 654, 2024. [Online]. Available: https://www.nature.com/articles/ s41467-024-44824-z
2024
-
[5]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in Proceedings of the International Conference on Machine Learning (ICML) , 2021, pp. 8748–8763
2021
-
[6]
Scaling up visual and vision-language representation learning with noisy text supervision,
C. Jia, Y . Yang, Y . Xia, Y .-T. Chen, Z. Parekh, H. Pham, Q. Le, Y .-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in International Conference on Machine Learning (ICML) , 2021, pp. 4904–4916
2021
-
[7]
BLIP: Bootstrapping Language- image Pre-training for Unified Vision-language Understanding and Generation,
J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrapping Language- image Pre-training for Unified Vision-language Understanding and Generation,” in International Conference on Machine Learning (ICML), 2022, pp. 12 888–12 900
2022
-
[8]
BLIP-2: Bootstrapping Language- image Pre-training With Frozen Image Encoders and Large Language Models,
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping Language- image Pre-training With Frozen Image Encoders and Large Language Models,” in International Conference on Machine Learning (ICML) , 2023, pp. 19 730–19 742
2023
Show all 107 references
-
[9]
Flamingo: A visual language model for few-shot learning,
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: A visual language model for few-shot learning,” Advances in Neural Information Processing Systems (NeurIPS), vol. 35, pp. 23 716–23 736, 2022
2022
-
[10]
Improved Baselines With Visual Instruction Tuning,
H. Liu, C. Li, Y . Li, and Y . J. Lee, “Improved Baselines With Visual Instruction Tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024, pp. 26 296– 26 306. 19
2024
-
[11]
Deepseek-vl: towards real-world vision-language understanding,
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang et al., “Deepseek-vl: towards real-world vision-language understanding,” arXiv preprint arXiv:2403.05525 , 2024
2024 arXiv
-
[12]
Microsoft COCO: Common Objects in Context,
T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ar, and C. L. Zitnick, “Microsoft COCO: Common Objects in Context,” in European Conference on Computer Vision (ECCV) , 2014, pp. 740–755
2014
-
[13]
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering,
Y . Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6904–6913
2017
-
[14]
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowl- edge,
K. Marino, M. Rastegari, A. Farhadi, and H. Hajishirzi, “OK-VQA: A Visual Question Answering Benchmark Requiring External Knowl- edge,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 3195–3204
2019
-
[15]
VizWiz Grand Challenge: Answering Visual Questions from Blind People,
D. Gurari, Q. Li, A. Stangl, C. Guo, C. Lin, K. Grauman, J. P. Bigham, and J. Luo, “VizWiz Grand Challenge: Answering Visual Questions from Blind People,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 3608– 3617
2018
-
[16]
Glam: Efficient Scaling of Lan- guage Models With Mixture-of-experts,
N. Du, Y . Huang, A. M. Dai, S. Tong, D. Lepikhin, Y . Xu, M. Krikun, Y . Zhou, A. W. Yu, O. Firat et al. , “Glam: Efficient Scaling of Lan- guage Models With Mixture-of-experts,” in International Conference on Machine Learning (ICML) , 2022, pp. 5547–5569
2022
-
[17]
Deepseek-vl2: Mixture-of-experts vision- language models for advanced multimodal understanding,
Z. Wu, X. Chen, Z. Pan, X. Liu, W. Liu, D. Dai, H. Gao, Y . Ma, C. Wu, B. Wang et al. , “Deepseek-vl2: Mixture-of-experts vision- language models for advanced multimodal understanding,” arXiv preprint arXiv:2412.10302, 2024
2024 arXiv
-
[18]
Chain-of-thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V . Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS), 2022
2022
-
[19]
Graph of thoughts: Solving elaborate problems with large language models,
M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk, and T. Hoefler, “Graph of thoughts: Solving elaborate problems with large language models,” Proceedings of the AAAI Conference on Artificial Intell...
2024 doi
-
[20]
Self-consistency improves chain of thought reasoning in language models,
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” arXiv preprint arXiv:2203.11171 , 2023. [Online]. Available: https://arxiv.org/abs/2203.11171
2023 arXiv
-
[21]
Retrieval- augmented Generation for Knowledge-intensive NLP Tasks,
P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. K ¨uttler, M. Lewis, W.-t. Yih, T. Rockt ¨aschel et al. , “Retrieval- augmented Generation for Knowledge-intensive NLP Tasks,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 33, pp. 9459...
2020
-
[22]
Language Mod- els are Few-shot Learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language Mod- els are Few-shot Learners,” Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 1877–1901, 2020
1901
-
[23]
Multitask prompted training enables zero-shot task generalization,
V . Sanh, A. Webson, C. Raffel, S. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, A. Raja, M. Dey, M. S. Bari, and Others, “Multitask prompted training enables zero-shot task generalization,” in International Conference on Learning Representations (ICLR) , 2022. [Onl...
2022
-
[24]
Scaling Instruction-finetuned Language Models,
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y . Tay, W. Fedus, Y . Li, X. Wang, M. Dehghani, S. Brahmaet al., “Scaling Instruction-finetuned Language Models,” Journal of Machine Learning Research (JMLR) , vol. 25, no. 70, pp. 1–53, 2024
2024
-
[25]
Chameleon: Mixed-modal early-fusion foundation models,
Chameleon Team, “Chameleon: Mixed-modal early-fusion foundation models,” arXiv preprint arXiv:2405.09818 , 2024. [Online]. Available: https://github.com/facebookresearch/chameleon
2024 arXiv
-
[26]
Training Language Models to Follow Instructions With Human Feedback,
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training Language Models to Follow Instructions With Human Feedback,” Advances in Neural Information Processing Systems (NeurIPS), vol. 35, pp. 27 730– 27 744, 2022
2022
-
[27]
Lora: Low-rank Adaptation of Large Language Models
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, W. Chen et al. , “Lora: Low-rank Adaptation of Large Language Models.” ICLR, vol. 1, no. 2, p. 3, 2022
2022
-
[28]
The Power of Scale for Parameter-efficient Prompt Tuning,
B. Lester, R. Al-Rfou, and N. Constant, “The Power of Scale for Parameter-efficient Prompt Tuning,” arXiv preprint arXiv:2104.08691 , 2021
2021 arXiv
-
[29]
Iris-SAM: Iris segmentation using a foundation model,
P. Farmanifard and A. Ross, “Iris-SAM: Iris segmentation using a foundation model,” in International Conference on Pattern Recognition and Artificial Intelligence , 2024, pp. 394–409
2024
-
[30]
ChatGPT meets iris biometrics,
——, “ChatGPT meets iris biometrics,” in IEEE International Joint Conference on Biometrics (IJCB) , 2024, pp. 1–10
2024
-
[31]
Dinov2: Learning Robust Visual Features Without Supervision,
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khali- dov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al., “Dinov2: Learning Robust Visual Features Without Supervision,” arXiv preprint arXiv:2304.07193, 2023
2023 arXiv
-
[32]
Froundation: Are foundation models ready for face recognition?
T. Chettaoui, N. Damer, and F. Boutros, “Froundation: Are foundation models ready for face recognition?” Image and Vision Computing, vol. 156, p. 105453, 2025
2025
-
[33]
A comprehensive survey of foundation models in medicine,
W. Khan, S. Leem, K. B. See, J. K. Wong, S. Zhang, and R. Fang, “A comprehensive survey of foundation models in medicine,” URL https://arxiv. org/abs/2406.10729, 2024
2024 arXiv
-
[34]
Biomedical foundation model: A survey,
X. Liu, Y . Zhang, Y . Lu, C. Yin, X. Hu, X. Liu, L. Chen, S. Wang, A. Rodriguez, H. Yao et al., “Biomedical foundation model: A survey,” arXiv preprint arXiv:2503.02104 , 2025
2025
-
[35]
Foundational models in medical imaging: A compre- hensive survey and future vision,
B. Azad, R. Azad, S. Eskandari, A. Bozorgpour, A. Kazerouni, I. Rekik, and D. Merhof, “Foundational models in medical imaging: A compre- hensive survey and future vision,” arXiv preprint arXiv:2310.18689 , 2023
2023 arXiv
-
[36]
A comprehensive survey on pretrained foundation models: A history from BERT to ChatGPT
C. Zhou, Q. Li, C. Li, J. Yu, Y . Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He et al., “A comprehensive survey on pretrained foundation models: A history from BERT to ChatGPT.” arXiv preprint arXiv:2302.09419, 2023
2023 arXiv
-
[37]
Foundation models defining a new era in vision: a survey and outlook,
M. Awais, M. Naseer, S. Khan, R. M. Anwer, H. Cholakkal, M. Shah, M.-H. Yang, and F. S. Khan, “Foundation models defining a new era in vision: a survey and outlook,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2025
2025
-
[38]
Vision-language models for vision tasks: A survey,
J. Zhang, J. Huang, S. Jin, and S. Lu, “Vision-language models for vision tasks: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024
2024
-
[39]
Multimodal large language models: A survey,
J. Wu, W. Gan, Z. Chen, S. Wan, and P. S. Yu, “Multimodal large language models: A survey,” in 2023 IEEE International Conference on Big Data (BigData) . IEEE, 2023, pp. 2247–2256
2023
-
[40]
A survey of multimodel large language models,
Z. Liang, Y . Xu, Y . Hong, P. Shang, Q. Wang, Q. Fu, and K. Liu, “A survey of multimodel large language models,” in Proceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineering, 2024, pp. 405–409
2024
-
[41]
Exploring the frontier of vision-language models: A survey of current methodologies and future directions,
A. Ghosh, A. Acharya, S. Saha, V . Jain, and A. Chadha, “Exploring the frontier of vision-language models: A survey of current methodologies and future directions,” arXiv preprint arXiv:2404.07214 , 2024
2024
-
[42]
Unleashing the potential of prompt engineering for large language models,
B. Chen, Z. Zhang, N. Langren ´e, and S. Zhu, “Unleashing the potential of prompt engineering for large language models,” Patterns, 2025
2025
-
[43]
Facex- former: A unified transformer for facial analysis,
K. Narayan, V . VS, R. Chellappa, and V . M. Patel, “Facex- former: A unified transformer for facial analysis,” arXiv preprint arXiv:2403.12960, 2024
2024 arXiv
-
[44]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning...
2021
-
[45]
Foundation models and biometrics: A survey and outlook,
H. O. Shahreza and S. Marcel, “Foundation models and biometrics: A survey and outlook,” TechRxiv, 2025
2025
-
[46]
Exploring ChatGPT for Face Presentation Attack Detection in Zero and Few-Shot in-Context Learning ,
A. Komaty, H. O. Shahreza, A. George, and S. Marcel, “ Exploring ChatGPT for Face Presentation Attack Detection in Zero and Few-Shot in-Context Learning ,” in IEEE/CVF Winter Conference on Applica- tions of Computer Vision Workshops (WACVW), 2025, pp. 1602–1611
2025
-
[47]
Towards zero-shot differential morphing attack detection with multimodal large language models,
R. Shekhawat, H. Li, R. Ramachandra, and S. Venkatesh, “Towards zero-shot differential morphing attack detection with multimodal large language models,” arXiv preprint arXiv:2505.15332 , 2025
2025 arXiv
-
[48]
Gemini: A family of highly capable multimodal models,
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, and Others, “Gemini: A family of highly capable multimodal models,” 2025. [Online]. Available: https://arxiv.org/abs/2312.11805
2025 arXiv
-
[49]
Iris recognition with off-the-shelf CNN features: A deep learning perspective,
K. Nguyen, C. Fookes, A. Ross, and S. Sridharan, “Iris recognition with off-the-shelf CNN features: A deep learning perspective,” IEEE Access, vol. 6, pp. 18 848–18 855, 2017
2017
-
[50]
Grother, M
P. Grother, M. Ngan, and K. Hanaoka, Face Recognition Vendor Test (FRVT): Part 3, Demographic Effects. National Institute of Standards and Technology Gaithersburg, MD, 2019
2019
-
[51]
Saving face: Investigating the ethical concerns of facial recognition auditing,
I. D. Raji, T. Gebru, M. Mitchell, J. Buolamwini, J. Lee, and E. Denton, “Saving face: Investigating the ethical concerns of facial recognition auditing,” in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 2020, pp. 145–151. 20
2020
-
[52]
Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification,
J. Buolamwini and T. Gebru, “Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification,” in Proceedings of the 1st Conference on Fairness, Accountability and Transparency , S. A. Friedler and C. Wilson, Eds., vol. 81, 23–24 Feb 2018, pp. 77–91. [On...
2018
-
[53]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
-
[54]
ImageNet Large Scale Visual Recognition Challenge,
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision , vol. 115, no. 3, pp. 211– 252, 2015
2015
-
[55]
EfficientNet: Rethinking model scaling for convolutional neural networks,
M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proceedings of the 36th International Conference on Machine Learning (ICML) , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97, 09–15...
2019
-
[56]
BERT: Pre- training of Deep Bidirectional Transformers for Language Understand- ing,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre- training of Deep Bidirectional Transformers for Language Understand- ing,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technol...
2019
-
[57]
Conceptual Cap- tions: A Cleaned, Hypernymed, Image Alt-text Dataset for Automatic Image Captioning,
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual Cap- tions: A Cleaned, Hypernymed, Image Alt-text Dataset for Automatic Image Captioning,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp....
2018
-
[58]
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hocken- maier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) . IEE...
2015
-
[59]
Reproducible scaling laws for contrastive language-image learning,
M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev, “Reproducible scaling laws for contrastive language-image learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR),...
2023
-
[60]
LAION- 400M: Open dataset of CLIP-filtered 400 million image-text pairs,
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki, “LAION- 400M: Open dataset of CLIP-filtered 400 million image-text pairs,” arXiv preprint arXiv:2111.02114 , 2021. [Online]. Available: https://arxiv.org/abs/2111.02114
2021 arXiv
-
[61]
Hugging Face: The AI Community Building the Future,
Hugging Face, “Hugging Face: The AI Community Building the Future,” https://huggingface.co, 2024, accessed: 2024-05-19
2024
-
[62]
LAION-5B: An open large- scale dataset for training next generation image-text models,
C. Schuhmann, R. Beaumont, R. Vencu, C. W. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. R. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev, “LAION-5B: An open large- scale dataset for training next generation...
2022
-
[63]
Emerging properties in self-supervised vision trans- formers,
M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision trans- formers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2021, pp. 9650–9660
2021
-
[64]
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,
W.-L. Chiang, Z. Li, Z. Lin, Y . Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y . Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,” March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/
2023
-
[65]
LLaMA: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “LLaMA: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[66]
Kosmos-2: Grounding multimodal large language models to the world,
Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, and F. Wei, “Kosmos-2: Grounding multimodal large language models to the world,” arXiv preprint arXiv:2306.14824 , 2023
2023 arXiv
-
[67]
InternVL3: exploring advanced training and test- time recipes for open-source multimodal models, (Technical Report),
J. Zhu, W. Wang, Z. Chen, Z. Liu, S. Ye, L. Gu, Y . Duan, H. Tian, W. Su, J. Shao et al., “InternVL3: exploring advanced training and test- time recipes for open-source multimodal models, (Technical Report),” arXiv preprint arXiv:2504.10479 , 2025
2025 arXiv
-
[68]
GPT-4o System Card,
OpenAI, “GPT-4o System Card,” https://openai.com/index/ gpt-4o-system-card/, 2024, accessed: 2025-05-11
2024
-
[69]
Claude 3.5 sonnet model card addendum,
Anthropic, “Claude 3.5 sonnet model card addendum,” https://www. anthropic.com/news/claude-3-5-sonnet, 2024, accessed: 2025-05-11
2024
-
[70]
Gemini 2.5: Our most intelligent AI model,
Google DeepMind, “Gemini 2.5: Our most intelligent AI model,” https://blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/, 2025, accessed: 2025- 05-11
2025
-
[71]
Deepseekmoe: Towards ultimate expert spe- cialization in mixture-of-experts language models,
D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y . Wu et al. , “Deepseekmoe: Towards ultimate expert spe- cialization in mixture-of-experts language models,” arXiv preprint arXiv:2401.06066, 2024
2024 arXiv
-
[72]
AgeDB: The first manually collected, in-the-wild age database,
S. Moschoglou, A. Papaioannou, C. Sagonas, J. Deng, I. Kotsia, and S. Zafeiriou, “AgeDB: The first manually collected, in-the-wild age database,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , 2017, pp. 51–59
2017
-
[73]
Nearest neighbor pattern classification,
T. M. Cover and P. E. Hart, “Nearest neighbor pattern classification,” IEEE Transactions on Information Theory , vol. 13, no. 1, pp. 21–27, 1967
1967
-
[74]
The use of multiple measurements in taxonomic prob- lems,
R. A. Fisher, “The use of multiple measurements in taxonomic prob- lems,” Annals of Eugenics , vol. 7, no. 2, pp. 179–188, 1936
1936
-
[75]
The regression analysis of binary sequences,
D. R. Cox, “The regression analysis of binary sequences,” Journal of the Royal Statistical Society: Series B (Methodological) , vol. 20, no. 2, pp. 215–242, 1958
1958
-
[76]
Ridge regression: Biased estimation for nonorthogonal problems,
A. E. Hoerl and R. W. Kennard, “Ridge regression: Biased estimation for nonorthogonal problems,” Technometrics, vol. 12, no. 1, pp. 55–67, 1970
1970
-
[77]
Support-vector networks,
C. Cortes and V . Vapnik, “Support-vector networks,” Machine Learn- ing, vol. 20, no. 3, pp. 273–297, 1995
1995
-
[78]
Labeled Faces in the Wild: A database for studying face recognition in uncon- strained environments,
G. B. Huang, M. Mattar, T. Berg, and E. Learned-Miller, “Labeled Faces in the Wild: A database for studying face recognition in uncon- strained environments,” in Workshop on Faces in ‘Real-Life’ Images: Detection, Alignment, and Recognition , 2008
2008
-
[79]
Cross-Pose LFW: A database for studying cross-pose face recognition in unconstrained environments,
T. Zheng and W. Deng, “Cross-Pose LFW: A database for studying cross-pose face recognition in unconstrained environments,” Beijing University of Posts and Telecommunications, Tech. Rep. 18-01, Febru- ary 2018
2018
-
[80]
Frontal to profile face verification in the wild,
S. Sengupta, J.-C. Chen, C. Castillo, V . M. Patel, R. Chellappa, and D. W. Jacobs, “Frontal to profile face verification in the wild,” in Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV), 2016, pp. 1–9
2016
-
[81]
Benchmarking Deep Network Architectures for Ethnicity Recognition Using a New Large Face Dataset,
A. Greco, G. Percannella, M. Vento, and V . Vigilante, “Benchmarking Deep Network Architectures for Ethnicity Recognition Using a New Large Face Dataset,” Machine Vision and Applications , vol. 31, pp. 1–13, 2020
2020
-
[82]
VGGFace2: A dataset for recognising faces across pose and age,
Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zisserman, “VGGFace2: A dataset for recognising faces across pose and age,” in 13th IEEE International Conference on Automatic Face & Gesture Recognition , 2018, pp. 67–74
2018
-
[83]
Comparison and Combination of Iris Matchers for Reliable Personal Authentication,
A. Kumar and A. Passi, “Comparison and Combination of Iris Matchers for Reliable Personal Authentication,” Pattern Recognition , vol. 43, no. 3, pp. 1016–1026, 2010
2010
-
[84]
Image Under- standing for Iris Biometrics: A survey,
K. W. Bowyer, K. Hollingsworth, and P. J. Flynn, “Image Under- standing for Iris Biometrics: A survey,” Computer Vision and Image Understanding, vol. 110, no. 2, pp. 281–307, 2008
2008
-
[85]
Ghclnet: A generalized hierarchically tuned contact lens detection network,
A. Singh, V . Mistry, D. Yadav, and A. Nigam, “Ghclnet: A generalized hierarchically tuned contact lens detection network,” in IEEE 4th International Conference on Identity, Security, and Behavior Analysis (ISBA), 2018, pp. 1–8
2018
-
[86]
IIT Delhi Iris Database (Ver- sion 1.0),
IIT Delhi Biometrics Research Lab, “IIT Delhi Iris Database (Ver- sion 1.0),” http://web.iitd.ac.in/ ∼biometrics/Database Iris.htm, 2008, accessed: 2025-05-23
2008
-
[87]
FaceForensics++: Learning to detect manipulated facial images,
A. R ¨ossler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nießner, “FaceForensics++: Learning to detect manipulated facial images,” in International Conference on Computer Vision (ICCV) , 2019
2019
-
[88]
Deepfakes,
GitHub, “Deepfakes,” https://github.com/deepfakes/faceswap, 2020, ac- cessed: 2025-05-15
2020
-
[89]
Face2Face: real-time face capture and reenactment of RGB videos,
J. Thies, M. Zollh ¨ofer, M. Stamminger, C. Theobalt, and M. Nießner, “Face2Face: real-time face capture and reenactment of RGB videos,” Communications of ACM , vol. 62, no. 1, p. 96–104, 2018. [Online]. Available: https://doi.org/10.1145/3292039
2018 doi
-
[90]
Faceswap,
M. GitHub, “Faceswap,” https://github.com/MarekKowalski/ FaceSwap/, 2021, accessed: 2025-05-15
2021
-
[91]
Deferred neural rendering: image synthesis using neural textures,
J. Thies, M. Zollh ¨ofer, and M. Nießner, “Deferred neural rendering: image synthesis using neural textures,” ACM Transactions of Graphics, vol. 38, no. 4, 2019. 21
2019
-
[92]
Extended StirTrace Benchmarking of Biometric and Forensic Qualities of Morphed Face Images,
T. Neubert, A. Makrushin, M. Hildebrandt, C. Kraetzer, and J. Dittmann, “Extended StirTrace Benchmarking of Biometric and Forensic Qualities of Morphed Face Images,” IET Biometrics , vol. 7, pp. 325–332, 2018
2018
-
[93]
Face Research Lab London (FRLL) Image Dataset,
L. DeBruine and B. Jones, “Face Research Lab London (FRLL) Image Dataset,” May 2017, https://figshare.com/articles/dataset/Face Research Lab London Set/5047666/3
2017
-
[94]
MorDIFF: Recognition Vulnerability and Attack Detectability of Face Morphing Attacks Created by Diffusion Autoencoders,
N. Damer, M. Fang, P. Siebke, J. N. Kolf, M. Huber, and F. Boutros, “MorDIFF: Recognition Vulnerability and Attack Detectability of Face Morphing Attacks Created by Diffusion Autoencoders,” in Proceedings of 11th International Workshop on Biometrics and Forensics (IWBF) , 2023...
2023
-
[95]
Face morph using OpenCV — C++ / Python,
S. Mallick, “Face morph using OpenCV — C++ / Python,” LearnOpenCV, 2016, https://learnopencv.com/ face-morph-using-opencv-cpp-python/
2016
-
[96]
Analyzing and Improving the Image Quality of StyleGAN,
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and Improving the Image Quality of StyleGAN,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
2020
-
[97]
debruine/webmorph morphing software: Beta release 2,
L. DeBruine, “debruine/webmorph morphing software: Beta release 2,” Jan. 2018, https://doi.org/10.5281/zenodo.1162670
2018 doi
-
[98]
Face morpher,
A. Quek, “Face morpher,” https://github.com/alyssaq/face-morpher, 2016, accessed: 2025-05-19
2016
-
[99]
dc-GAN: Dual-Conditioned GAN for Face Demorphing From a Single Morph,
N. Shukla and A. Ross, “dc-GAN: Dual-Conditioned GAN for Face Demorphing From a Single Morph,” in Proceedings of IEEE Interna- tional Conference on Automatic Face and Gesture Recognition , 2025
2025
-
[100]
Metric for Evaluating Performance of Reference-Free Demor- phing Methods,
——, “Metric for Evaluating Performance of Reference-Free Demor- phing Methods,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW) , 2025, pp. 1585–1591
2025
-
[101]
Privacy-Friendly Synthetic Data for the Development of Face Morphing Attack Detectors,
N. Damer, C. A. F. L ´opez, M. Fang, N. Spiller, M. V . Pham, and F. Boutros, “Privacy-Friendly Synthetic Data for the Development of Face Morphing Attack Detectors,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 20...
2022
-
[102]
LLaV A-NeXT: Improved reasoning, ocr, and world knowledge,
H. Liu, C. Li, Y . Li, B. Li, Y . Zhang, S. Shen, and Y . J. Lee, “LLaV A-NeXT: Improved reasoning, ocr, and world knowledge,” January 2024. [Online]. Available: https://llava-vl.github.io/blog/ 2024-01-30-llava-next/
2024
-
[103]
CosFace: Large margin cosine loss for deep face recognition,
H. Wang, Y . Wang, Z. Zhou, X. Ji, D. Gong, J. Zhou, Z. Li, and W. Liu, “CosFace: Large margin cosine loss for deep face recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 5265–5274
2018
-
[104]
ArcFace: Additive angular margin loss for deep face recognition,
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “ArcFace: Additive angular margin loss for deep face recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 4690–4699
2019
-
[105]
AdaFace: Quality adaptive margin for face recognition,
M. Kim, A. K. Jain, and X. Liu, “AdaFace: Quality adaptive margin for face recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022
2022
-
[106]
MagFace: A universal representation for face recognition and quality assessment,
Q. Meng, S. Zhao, Z. Huang, and F. Zhou, “MagFace: A universal representation for face recognition and quality assessment,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 14 225–14 234
2021
-
[107]
The birth of self supervised learning: A supervised theory,
R. Balestriero and Y . LeCun, “The birth of self supervised learning: A supervised theory,” in NeurIPS 2024 Workshop: Self- Supervised Learning - Theory and Practice , 2024. [Online]. Available: https://openreview.net/forum?id=NhY AjAAdQT 22
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.