Pith. sign in

REVIEW 5 cited by

Benchmarking Foundation Models for Zero-Shot Biometric Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.24214 v1 pith:7NC6CQH4 submitted 2025-05-30 cs.CV cs.AI

Benchmarking Foundation Models for Zero-Shot Biometric Tasks

classification cs.CV cs.AI
keywords biometricmodelsfacedetectionpercenttasksfoundationiris
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The advent of foundation models, particularly Vision-Language Models (VLMs) and Multi-modal Large Language Models (MLLMs), has redefined the frontiers of artificial intelligence, enabling remarkable generalization across diverse tasks with minimal or no supervision. Yet, their potential in biometric recognition and analysis remains relatively underexplored. In this work, we introduce a comprehensive benchmark that evaluates the zero-shot and few-shot performance of state-of-the-art publicly available VLMs and MLLMs across six biometric tasks spanning the face and iris modalities: face verification, soft biometric attribute prediction (gender and race), iris recognition, presentation attack detection (PAD), and face manipulation detection (morphs and deepfakes). A total of 41 VLMs were used in this evaluation. Experiments show that embeddings from these foundation models can be used for diverse biometric tasks with varying degrees of success. For example, in the case of face verification, a True Match Rate (TMR) of 96.77 percent was obtained at a False Match Rate (FMR) of 1 percent on the Labeled Face in the Wild (LFW) dataset, without any fine-tuning. In the case of iris recognition, the TMR at 1 percent FMR on the IITD-R-Full dataset was 97.55 percent without any fine-tuning. Further, we show that applying a simple classifier head to these embeddings can help perform DeepFake detection for faces, Presentation Attack Detection (PAD) for irides, and extract soft biometric attributes like gender and ethnicity from faces with reasonably high accuracy. This work reiterates the potential of pretrained models in achieving the long-term vision of Artificial General Intelligence.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VISER: Visually-Informed System for Enhanced Robustness in Open-Set Iris Presentation Attack Detection

    cs.CV 2026-03 unverdicted novelty 6.0

    Denoised eye tracking heatmaps yield the largest generalization gain in open-set iris presentation attack detection, lowering APCER at 1% BPCER compared to cross-entropy training.

  2. FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis

    cs.CV 2025-12 conditional novelty 6.0

    FPBench evaluates 20 MLLMs across 8 fingerprint tasks on 7 datasets and shows fine-tuning vision and language encoders improves performance by 7-39%.

  3. A Systematic Failure Analysis of Vision Foundation Models for Open Set Iris Presentation Attack Detection

    cs.CV 2026-05 accept novelty 5.0

    Vision foundation models transfer across similar iris datasets but fail to generalize to unseen presentation attacks and cross-spectral shifts in open-set PAD.

  4. Are Face Embeddings Compatible Across Deep Neural Network Models?

    cs.CV 2026-04 unverdicted novelty 5.0

    Simple affine transformations align face embeddings across different DNN models, substantially improving cross-model identification and verification performance.

  5. DiscoGen: Procedural Generation of Algorithm Discovery Tasks in Machine Learning

    cs.LG 2026-03 conditional novelty 5.0

    DiscoGen procedurally generates billions of configurable ML algorithm-discovery tasks and a fixed DiscoBench subset so algorithm-discovery agents can be trained and evaluated without contamination or saturation.