Pith. sign in

REVIEW 4 major objections 6 minor 34 references

On the Use of Synthetic Data for Threshold Calibration in Face Recognition: Performance and Security Implications for Border Control Systems

T0 review · 4 major / 6 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read Synthetic face data can set preliminary verification thresholds, but low-FMR border systems still need real operational data because impostor score tails do not transfer.

desk verdict Solid empirical negative result on synthetic low-FMR threshold transfer, but the “tail mismatch” story is underspecified and likely confounded by bulk score shift. read the letter →

arxiv 2607.25990 v1 pith:WWXCH5OO submitted 2026-07-27 cs.CV

classification cs.CV
keywords facerecognitionthresholdcalibrationsyntheticdatalowFMRbordercontrolEntry/ExitSystemmorphattackscoredistributiontails
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

European border systems such as the Entry/Exit System must run face verification at extremely low false-match rates, yet Member States often lack legal and practical access to representative real images for setting the accept/reject threshold. This paper tests whether synthetic document-to-live face pairs can supply those thresholds instead. It shows that synthetic data can roughly match calibration behavior under controlled capture, but mismatches in the extreme tails of the impostor score distribution cause thresholds to fail under unconstrained conditions, raising false non-matches and leaving systems more open to morph attacks. Even different synthetic generators disagree with each other. The practical conclusion is that synthetic data is useful for development and first-pass calibration, while final high-security thresholds still require validation and adjustment on real operational data.

What carries the argument

Cross-domain transfer of decision thresholds taken as extreme quantiles of the impostor cosine-similarity distribution (target FMRs from 0.1% down to 0.001%), measured by how well those thresholds preserve target FMR and acceptable FNMR, plus morph nearer/farther match passing rates, when moved from synthetic calibration sets to real or other synthetic evaluation sets.

What would settle it

Re-run the same quantile-threshold transfer protocol with a higher-capacity recognition model and with operational EES-style capture streams (kiosk, guided booth, and mobile) that include sensor and demographic variation; if synthetic-calibrated thresholds then hold target FMR and FNMR at 0.001% without large morph-acceptance spikes, the central claim is weakened.

Watch

Extended reading notes

Core claim

Synthetic face datasets can approximate threshold calibration for document-to-live face verification in controlled settings, but they do not reliably transfer to unconstrained conditions at the low false-match rates required by border control; the failures arise from mismatches in impostor score-distribution tails and produce both large rises in false non-match rate and higher acceptance of morph attacks, so reliable operational thresholds still need real-world validation.

Load-bearing premise

Two real frontal datasets (one tightly controlled, one unconstrained) plus a single lightweight face model are treated as enough to stand in for the full range of real border capture conditions.

Editorial extensions

If this is right

  • Member-State border systems cannot treat synthetic-only calibration as final for EES-grade low-FMR thresholds.
  • Preliminary synthetic thresholds still need explicit safety margins and a real-data adjustment step before go-live.
  • Morph-attack testing must be part of any calibration pipeline, because permissive synthetic thresholds raise both nearer- and farther-match acceptance.
  • Even among synthetic generators, score-distribution differences are large enough that the choice of generator itself changes the resulting operating point.
  • Future synthetic generators aimed at calibration must deliberately match operational impostor tails, not only average genuine/impostor separation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Privacy and legal barriers to real biometric data will keep pressure on synthetic calibration, so hybrid pipelines that mix limited real operational samples with large synthetic sets are the likely near-term engineering path.
  • The same tail-sensitivity problem should appear in any biometric modality (fingerprint, iris) that must set ultra-low FMR thresholds from non-operational data.
  • Regulators writing EES-style performance targets may eventually need to require documented real-data threshold validation rather than leaving calibration method unspecified.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The manuscript studies whether synthetic face datasets can be used to calibrate document-to-live verification thresholds for EES-like border control. Using EdgeFace-S, FLUXSyn-ID and Syn-Multi-PIE (plus a combined synthetic set) for calibration, and Color FERET frontal and CFPW frontal as controlled/unconstrained real proxies, the authors estimate raw cosine thresholds as impostor quantiles at target FMRs from 0.1% to 0.001% and measure cross-dataset FNMR/observed FMR and AMSL morph NMPR/FMPR. They report that FLUXSyn-ID transfers reasonably to FERET but all synthetic calibrations degrade sharply on CFPW, that synthetic-to-synthetic transfer is asymmetric, and that looser thresholds increase morph acceptance. The conclusion is that synthetic data are useful for development/preliminary calibration but reliable low-FMR threshold selection usually requires representative real-data validation.

Significance. If the central claim holds, the work is practically significant for EES-like deployments: it provides a clean no-image-overlap cross-dataset protocol, focuses on the difficult 0.1%-0.001% FMR regime rather than EER, shows quantitatively that thresholds are dataset-dependent even among synthetic sets, and connects calibration errors to morph-attack exposure. The use of an open lightweight EdgeFace model, named public/prior datasets, explicit operating-point grid, and directly falsifiable transfer tables are strengths. The paper does not appear circular: thresholds are estimated on held-out impostor scores and then tested on disjoint datasets. Its value would increase substantially if the mechanism for transfer failure were isolated and uncertainty at extreme quantiles were quantified.

major comments (4)
  1. [Abstract, §3.3, §5.1, Tables 1,3,5] Abstract, §3.3, §5.1, Tables 1/3/5: the central mechanism is stated as failure 'due to mismatches in score distribution tails', but the reported signature is also exactly what a bulk location/scale shift would produce. Impostor mean cosine differs by an order of magnitude across datasets (Table 1: FLUXSyn-ID 0.107, Syn-Multi-PIE 0.199, CFPW/FERET 0.015), while thresholds are raw-cosine impostor quantiles (§3.3; Table 3). Thus a FLUXSyn 0.001% threshold of 0.630 sits far above essentially all CFPW impostor scores; Table 5 then shows observed FMR collapsing to <0.0002% while FNMR reaches 25.79%. FMR far below target plus inflated FNMR is the fingerprint of global threshold strictness/unit mismatch, not specifically tail-shape error. Please isolate the mechanism: report impostor mean/std and quantile ratios; re-run transfer after impostor z-normalization or an affine map estimated from the
  2. [§3.4, §5.3, Table 8] Table 8 and §5.3: the morph-security claim compares NMPR/FMPR at nominal target FMRs, but the calibrations do not achieve comparable operating points on real data (Tables 4-5). High FLUXSyn NMPR/FMPR at 0.10% (0.997/0.734) versus lower Syn-Multi-PIE values may simply reflect more permissive absolute thresholds rather than a synthetic-calibration-specific vulnerability. The load-bearing comparison should be iso-FMR: choose thresholds that achieve the same observed FMR on a held-out reference real set (or same FNMR), then report NMPR/FMPR, ideally with the morph set's genuine FNMR and a real-data-calibrated baseline. Without this, 'increased vulnerability to morph-based attacks' in the abstract is broader than the experiment supports.
  3. [§3.3, Tables 4-7] Tables 4-7, §3.3: extreme quantile estimates are used without uncertainty. At FMR=0.001% (1e-5), Color FERET's 3.7M impostor comparisons imply only ~37 expected false matches, and many observed FMR cells are censored as '<0.0002%'; FNMR=0.000% on FERET is also compatible with zero errors among only 3,652 genuine items. The strong sensitivity claimed at 0.005%/0.001% could be partly estimation noise. Add bootstrap or Wilson/binomial confidence intervals for FNMR/FMR and calibrated quantiles, state effective sample sizes/minimum detectable rates, and mark cells whose uncertainty crosses the target FMR.
  4. [§3.1, Table 1] Table 1 and §3.1 pair generation: the table is ambiguous about whether 'Gen.' and 'Imp.' are images or pairs. FLUXSyn-ID 44,653 and Syn-Multi-PIE 30,000 look like image counts (approx. IDs × 3 types), whereas the text repeatedly defines genuine pairs as all within-identity comparisons; CFPW is listed with 21,735 'Gen.' despite 500 identities×20 images implying a different pair count. Since impostor pair counts determine quantile resolution and the censored FMR cells, please give exact genuine/impostor pair counts, sampling/exhaustion rules, deduplication, identity-disjointness, and random seeds, and reconcile Table 1 with the protocol.
minor comments (6)
  1. [Figs. 4-5, §3.3] Figures 4-5 are called 'distance distributions' but the protocol uses cosine similarity; use one term consistently and state explicitly that higher cosine = stricter acceptance threshold.
  2. [§3.1] §3.1 repeats the sentence defining genuine pairs for each dataset; consolidate once in Pair generation. Also define 'reference', 'illumination_0', 'pose_3', f_doc, live_0e_d1 and live_0p_d1 in a small table rather than inline fragments.
  3. [§3.2, Table 2] Table 2 mixes LFW accuracy from heterogeneous publications; add a caution that these are not directly comparable and are used only to motivate EdgeFace-S, not to rank models for this study.
  4. [§3.2-§3.3] Release the threshold-estimation/evaluation scripts, pair lists or hashes, EdgeFace checkpoint identifier, alignment pipeline, and seeds; the protocol is otherwise reproducible but the paper currently reports no artifact.
  5. [Abstract, §5.2, §7] In §5.2, phrase EES relevance as a proxy argument: FERET/CFPW frontal subsets plus one lightweight model bound controlled/unconstrained extremes but do not include kiosk/mobile sensors, quality filters, demographics, or multi-stage logic; the abstract should mirror this scoping.
  6. [Throughout] Typos/OCR artifacts in the arXiv text (e.g., 'Fac e', 'verificatio n', ligature breaks) should be cleaned before camera-ready; check that tables' '* <0.0002%' censoring is explained in captions.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: standard empirical cross-dataset quantile transfer; results do not reduce to inputs by construction.

full rationale

The paper’s load-bearing chain is purely empirical. Thresholds are defined as quantiles of impostor cosine-similarity scores on a calibration set (§3.3), then applied without image overlap to held-out evaluation sets; FNMR, observed FMR, NMPR, and FMPR are measured directly (Tables 3–8). Nothing is fitted and then re-labeled as a prediction, and no uniqueness theorem or ansatz is imported from the author’s prior work to force the conclusion. Synthetic datasets (FLUXSyn-ID, Syn-Multi-PIE), the EdgeFace model, Color FERET, CFPW, and AMSL morphs are external resources; citing them does not make the transfer failures tautological. The claim that synthetic calibration fails under domain shift because of score-distribution mismatch is an observational finding from those tables, not a restatement of the calibration definition. Correctness concerns (e.g., bulk location/scale shift vs. tail shape, lack of score normalization) are outside the scope of circularity. Score 0 is appropriate.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

Load-bearing content is experimental, not axiomatic derivation. The claim rests on standard biometric score-threshold definitions, the representativeness of chosen synthetic/real/morph corpora and one edge model, and the operational reading that FERET/CFPW-style extremes inform EES-like calibration policy.

free parameters (3)
  • Target FMR operating grid (0.1%, 0.05%, 0.01%, 0.005%, 0.001%) = 0.1% to 0.001%
    Hand-chosen quantiles reflecting border-control practice and older Frontex/EES figures; they define all reported thresholds and drive the ‘low-FMR sensitivity’ narrative.
  • Impostor pair sampling size / procedure = up to ~100M–200M impostor pairs (dataset-dependent)
    Table 1 reports very large impostor counts (e.g., 100M); random cross-identity sampling details and any caps affect extreme quantile estimates that set thresholds.
  • EdgeFace-S gamma=0.5 choice = EdgeFace-S γ=0.5 (3.65M params)
    Single architecture/scale selected for ‘realistic edge deployment’; all scores and thresholds are conditional on this embedding space.
assumptions (4)
  • domain assumption Decision thresholds at target FMR equal empirical quantiles of the impostor cosine-similarity distribution on the calibration set.
    Standard biometric operating-point definition used throughout §3.3 and Table 3; central transfer claims assume this calibration method.
  • domain assumption Cosine similarity on frozen EdgeFace embeddings is an adequate score for studying operational threshold transfer.
    §3.2 uses the model as provided without fine-tuning; no other scorers or calibration mappings are tested.
  • ad hoc to paper Color FERET frontal and CFPW frontal bound controlled vs unconstrained document-to-live conditions relevant to EES-like settings.
    §3.1 and Scope paragraph treat these as extremes for domain shift; operational recommendation depends on that proxy choice.
  • domain assumption Landmark-based AMSL morphs are a meaningful security probe for threshold miscalibration.
    §3.4 and Table 8; NMPR/FMPR under transferred thresholds support the security half of the claim.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the Use of Synthetic Data for Threshold Calibration in Face Recognition: Performance and Security Implications for Border Control Systems." pith.science (2026). https://pith.science/paper/WWXCH5OO

@misc{pith2026260725990,
  author       = {Pith},
  title        = {Pith review of: On the Use of Synthetic Data for Threshold Calibration in Face Recognition: Performance and Security Implications for Border Control Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WWXCH5OO}},
  note         = {Machine review of arXiv:2607.25990}
}
read the original abstract

The recently deployed Entry/Exit System (EES) introduces large-scale biometric verification into European border control, requiring face recognition systems to operate at extremely low false match rates (FMR). While regulatory frameworks define performance targets at the EES Central System level, they do not specify how verification thresholds should be calibrated in practice at the Member State level. In operational settings, obtaining representative real-world data for calibration is often constrained by legal, logistical, and privacy limitations. In this work, we investigate the use of synthetic face data for threshold calibration in document-to-live verification scenarios relevant to border control systems. We analyze the alignment of genuine and impostor score distributions between synthetic and real datasets and evaluate the transferability of calibrated thresholds across domains, with a focus on low-FMR operating points. Our results show that synthetic data can approximate calibration behavior in controlled settings, but fails to reliably generalize to unconstrained conditions due to mismatches in score distribution tails. These discrepancies lead to significant degradation in recognition performance and increased vulnerability to morph-based attacks. We further demonstrate that calibration outcomes are highly dataset-dependent, even across synthetic datasets. Overall, our findings highlight that while synthetic data is useful for system development and preliminary calibration, our results indicate that reliable threshold selection in high-security deployments typically requires validation and adjustment using representative real-world data.

Figures

Figures reproduced from arXiv: 2607.25990 by the authors.

Figure 1
Figure 1. Examples of synthetic images by dataset and image t [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Examples of Color FERET and CFPW frontal images. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Example AMSL morph images. Pair generation. For each dataset, genuine (mated) pairs are formed by comparing images of the same identity. Impostor (non-mated) pairs are generated by comparing im￾ages across different identities. Genuine pairs are formed from images of the same identity, while impostor pairs are randomly sampled across identities to obtain representative distributions. Dataset IDs Gen. Imp. Gen. x¯ Im… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Distance distributions on (a) Color FERET frontal [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Distance distributions and calculated threshold [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

34 extracted references · 5 linked inside Pith

  1. [1]

    G. Bae, M. de La Gorce, T. Baltruvaitis, C. Hewitt, D. Chen , J. V alentin, R. Cipolla, and J. Shen. Digiface-1m: 1 million digital face images for face recognition. In IEEE/CVF Win- ter Conference on Applications of Computer Vision (WACV), pages 3526–3535. IEEE, 2023

  2. [2]

    Borsukiewicz, F

    P . Borsukiewicz, F. Boutros, I. E. Olatunji, C. Beumier, W. C. Ouedraogo, J. Klein, and T. F. Bissyand´ e. Beyond real faces: Synthetic datasets can achieve reliable recognitio n performance without privacy compromise. arXiv preprint arXiv:2510.17372, 2025

  3. [3]

    Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zisserman. Vggface2: A dataset for recognising faces across pose and age. In 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG), pages 67–74. IEEE, 2018

  4. [4]

    S. Chen, Y . Liu, X. Gao, and Z. Han. Mobilefacenets: Effi- cient cnns for accurate real-time face verification on mobil e devices. In CCBR, 2018

  5. [5]

    Colbois, T

    L. Colbois, T. de Freitas Pereira, and S. Marcel. On the us e of automatically generated synthetic image datasets for benc h- marking face recognition. In IEEE International Joint Con- ference on Biometrics (IJCB) , pages 1–8. IEEE, 2021

  6. [6]

    DeBruine and B

    L. DeBruine and B. Jones. Face research lab london set, 2017

  7. [7]

    J. Deng, J. Guo, N. Xue, and S. Zafeiriou. Arcface: Ad- ditive angular margin loss for deep face recognition. In 2019 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR) , pages 4685–4694, 2019

  8. [8]

    Independent eval- uation of biometric technology in the eu: State of play and future options, 2026

    eu-LISA, European Commission DG HOME, European Commission JRC, Europol, and Frontex. Independent eval- uation of biometric technology in the eu: State of play and future options, 2026. Policy Brief. Available on- line: https://www.eulisa.europa.eu/our-publications/ policy- brief-independent-evaluation-biometric-technology-eu

Show all 34 references
  1. [9]

    European Commission. Commission implementing decisio n (eu) 2019/329 laying down the specifications for the quality , resolution and use of fingerprints and facial image for bio- metric verification and identification in the entry/exit sys tem (ees), February 2019. Last accessed: ...

  2. [10]

    Commission announces launch of shared biometric matching service, May 2025

    European Commission. Commission announces launch of shared biometric matching service, May 2025. Last ac- cessed: 23 Apr 2026

  3. [11]

    Entry/exit system (ees), April 2 026

    European Commission. Entry/exit system (ees), April 2 026. Last accessed: 23 Apr 2026

  4. [12]

    European Parliament and the Council of the European Union. Regulation (eu) 2017/2226 establishing an entry/ex it system (ees) to register entry and exit data and refusal of en - try data of third-country nationals crossing the external b or- ders of the member states, November...

  5. [13]

    Best practice technical guidelines for automated border contro l (abc) systems, August 2015

    Frontex - European Border and Coast Guard Agency. Best practice technical guidelines for automated border contro l (abc) systems, August 2015. Last accessed: 23 Apr 2026

  6. [14]

    Tech- nical guide for border checks on entry/exit system (ees) re- lated equipment, June 2021

    Frontex - European Border and Coast Guard Agency. Tech- nical guide for border checks on entry/exit system (ees) re- lated equipment, June 2021. Last accessed: 23 Apr 2026

  7. [15]

    George, C

    A. George, C. Ecabert, H. O. Shahreza, K. Kotwal, and S. Marcel. Edgeface: Efficient face recognition model for edge devices. IEEE Transactions on Biometrics, Behavior , and Identity Science, 2024

  8. [16]

    Granoviter, A

    O. Granoviter, A. Gruzdev, V . Loginov, M. Kogan, and O. Zvitia. Face recognition using synthetic face data. arXiv preprint arXiv:2305.10079, 2023

  9. [17]

    G. B. Huang, M. Mattar, T. Berg, and E. Learned-Miller. L a- beled faces in the wild: A database for studying face recog- nition in unconstrained environments. In W orkshop on Faces in ’Real-Life’ Images: Detection, Alignment, and Recogni- tion, 2008

  10. [18]

    Ismayilov, D

    R. Ismayilov, D. Sero, and L. Spreeuwers. Fluxsynid: A framework for identity-controlled synthetic face gen- eration with document and live images. arXiv preprint arXiv:2505.07530, 2025

  11. [19]

    A. K. Jain, A. Ross, and K. Nandakumar. Introduction to Biometrics. Springer, 2011

  12. [20]

    B. F. Klare and A. K. Jain. Heterogeneous face recogniti on: Matching nir to visible light images. In IJCB, 2010

  13. [21]

    Liu et al

    J. Liu et al. Oneface: One threshold for all. In Computer Vision – ECCV 2022 W orkshops , volume 13801 of Lecture Notes in Computer Science . Springer, 2022

  14. [22]

    Q. Meng, S. Zhao, Z. Huang, and F. Zhou. Magface: A universal representation for face recognition and quality as- sessment. In 2021 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR) , pages 14220–14229, 2021

  15. [23]

    Neubert, A

    T. Neubert, A. Makrushin, M. Hildebrandt, C. Kraetzer, and J. Dittmann. Extended stirtrace benchmarking of biometric and forensic qualities of morphed face images. IET Biomet- rics, 7(4):325–332, 2018

  16. [24]

    Nguyen, C

    K. Nguyen, C. Fookes, S. Sridharan, and S. Denman. Deep learning for face recognition: A survey. arXiv preprint arXiv:1804.06655, 2018

  17. [25]

    Color feret database

    NIST. Color feret database. https://www.nist.gov/itl/products-and-services/color-feret- database

  18. [26]

    Face recognition technology (feret) program

    NIST. Face recognition technology (feret) program. https://www.nist.gov/programs-projects/face-recognition- technology-feret

  19. [27]

    H. Qiu, B. Y u, D. Gong, Z. Li, W. Liu, and D. Tao. Syn- face: Face recognition with synthetic data. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 10860–10870. IEEE, 2021

  20. [28]

    Schroff, D

    F. Schroff, D. Kalenichenko, and J. Philbin. Facenet: A uni- fied embedding for face recognition and clustering. In 2015 IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 815–823, 2015

  21. [29]

    Sengupta, J

    S. Sengupta, J. C. Cheng, C. D. Castillo, V . M. Patel, R. Chellappa, and D. W. Jacobs. Frontal to profile face veri- fication in the wild. In IEEE Conference on Applications of Computer Vision, February 2016

  22. [30]

    H. O. Shahreza, C. Ecabert, A. George, A. Unnervik, S. Ma r- cel, N. D. Domenico, G. Borghi, D. Maltoni, F. Boutros, J. V ogel, N. Damer, ˜A. Sanchez-Perez, E. Mas-Candela, J. Calvo-Zaragoza, B. Biesseck, P . Vidal, R. Granada, D. Menotti, I. Deandres-Tame, and J. Fierrez. Sdf...

  23. [31]

    Y . Sun, D. Liang, X. Wang, and X. Tang. Deepid3: Face recognition with very deep neural networks. ArXiv, abs/1502.00873, 2015

  24. [32]

    Wu and K

    H. Wu and K. W. Bowyer. What should be balanced in a ”balanced” face recognition dataset? In Proceedings of the British Machine Vision Conference (BMVC) , page 235. BMV A Press, 2023

  25. [33]

    Yeung, T

    M. Yeung, T. Teramoto, S. Wu, T. Fujiwara, K. Suzuki, and T. Kojima. V ariface: Fair and diverse synthetic dataset gener- ation for face recognition. arXiv preprint arXiv:2412.06235, 2024

  26. [34]

    Z. Zhu, G. Huang, J. Deng, Y . Ye, J. Huang, X. Chen, J. Zhu, T. Yang, J. Lu, D. Du, and J. Zhou. Webface260m: A bench- mark unveiling the power of million-scale deep face recog- nition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 2021. IEEE

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.