REVIEW 4 major objections 6 minor 34 references
On the Use of Synthetic Data for Threshold Calibration in Face Recognition: Performance and Security Implications for Border Control Systems
T0 review · 4 major / 6 minor · reviewed 2026-07-31 · grok-4.5
Pith's one-line read Synthetic face data can set preliminary verification thresholds, but low-FMR border systems still need real operational data because impostor score tails do not transfer.
desk verdict Solid empirical negative result on synthetic low-FMR threshold transfer, but the “tail mismatch” story is underspecified and likely confounded by bulk score shift. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Cross-domain transfer of decision thresholds taken as extreme quantiles of the impostor cosine-similarity distribution (target FMRs from 0.1% down to 0.001%), measured by how well those thresholds preserve target FMR and acceptable FNMR, plus morph nearer/farther match passing rates, when moved from synthetic calibration sets to real or other synthetic evaluation sets.
What would settle it
Re-run the same quantile-threshold transfer protocol with a higher-capacity recognition model and with operational EES-style capture streams (kiosk, guided booth, and mobile) that include sensor and demographic variation; if synthetic-calibrated thresholds then hold target FMR and FNMR at 0.001% without large morph-acceptance spikes, the central claim is weakened.
Extended reading notes
Core claim
Synthetic face datasets can approximate threshold calibration for document-to-live face verification in controlled settings, but they do not reliably transfer to unconstrained conditions at the low false-match rates required by border control; the failures arise from mismatches in impostor score-distribution tails and produce both large rises in false non-match rate and higher acceptance of morph attacks, so reliable operational thresholds still need real-world validation.
Load-bearing premise
Two real frontal datasets (one tightly controlled, one unconstrained) plus a single lightweight face model are treated as enough to stand in for the full range of real border capture conditions.
Editorial extensions
If this is right
- Member-State border systems cannot treat synthetic-only calibration as final for EES-grade low-FMR thresholds.
- Preliminary synthetic thresholds still need explicit safety margins and a real-data adjustment step before go-live.
- Morph-attack testing must be part of any calibration pipeline, because permissive synthetic thresholds raise both nearer- and farther-match acceptance.
- Even among synthetic generators, score-distribution differences are large enough that the choice of generator itself changes the resulting operating point.
- Future synthetic generators aimed at calibration must deliberately match operational impostor tails, not only average genuine/impostor separation.
Reading between the lines
- Privacy and legal barriers to real biometric data will keep pressure on synthetic calibration, so hybrid pipelines that mix limited real operational samples with large synthetic sets are the likely near-term engineering path.
- The same tail-sensitivity problem should appear in any biometric modality (fingerprint, iris) that must set ultra-low FMR thresholds from non-operational data.
- Regulators writing EES-style performance targets may eventually need to require documented real-data threshold validation rather than leaving calibration method unspecified.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies whether synthetic face datasets can be used to calibrate document-to-live verification thresholds for EES-like border control. Using EdgeFace-S, FLUXSyn-ID and Syn-Multi-PIE (plus a combined synthetic set) for calibration, and Color FERET frontal and CFPW frontal as controlled/unconstrained real proxies, the authors estimate raw cosine thresholds as impostor quantiles at target FMRs from 0.1% to 0.001% and measure cross-dataset FNMR/observed FMR and AMSL morph NMPR/FMPR. They report that FLUXSyn-ID transfers reasonably to FERET but all synthetic calibrations degrade sharply on CFPW, that synthetic-to-synthetic transfer is asymmetric, and that looser thresholds increase morph acceptance. The conclusion is that synthetic data are useful for development/preliminary calibration but reliable low-FMR threshold selection usually requires representative real-data validation.
Significance. If the central claim holds, the work is practically significant for EES-like deployments: it provides a clean no-image-overlap cross-dataset protocol, focuses on the difficult 0.1%-0.001% FMR regime rather than EER, shows quantitatively that thresholds are dataset-dependent even among synthetic sets, and connects calibration errors to morph-attack exposure. The use of an open lightweight EdgeFace model, named public/prior datasets, explicit operating-point grid, and directly falsifiable transfer tables are strengths. The paper does not appear circular: thresholds are estimated on held-out impostor scores and then tested on disjoint datasets. Its value would increase substantially if the mechanism for transfer failure were isolated and uncertainty at extreme quantiles were quantified.
major comments (4)
- [Abstract, §3.3, §5.1, Tables 1,3,5] Abstract, §3.3, §5.1, Tables 1/3/5: the central mechanism is stated as failure 'due to mismatches in score distribution tails', but the reported signature is also exactly what a bulk location/scale shift would produce. Impostor mean cosine differs by an order of magnitude across datasets (Table 1: FLUXSyn-ID 0.107, Syn-Multi-PIE 0.199, CFPW/FERET 0.015), while thresholds are raw-cosine impostor quantiles (§3.3; Table 3). Thus a FLUXSyn 0.001% threshold of 0.630 sits far above essentially all CFPW impostor scores; Table 5 then shows observed FMR collapsing to <0.0002% while FNMR reaches 25.79%. FMR far below target plus inflated FNMR is the fingerprint of global threshold strictness/unit mismatch, not specifically tail-shape error. Please isolate the mechanism: report impostor mean/std and quantile ratios; re-run transfer after impostor z-normalization or an affine map estimated from the
- [§3.4, §5.3, Table 8] Table 8 and §5.3: the morph-security claim compares NMPR/FMPR at nominal target FMRs, but the calibrations do not achieve comparable operating points on real data (Tables 4-5). High FLUXSyn NMPR/FMPR at 0.10% (0.997/0.734) versus lower Syn-Multi-PIE values may simply reflect more permissive absolute thresholds rather than a synthetic-calibration-specific vulnerability. The load-bearing comparison should be iso-FMR: choose thresholds that achieve the same observed FMR on a held-out reference real set (or same FNMR), then report NMPR/FMPR, ideally with the morph set's genuine FNMR and a real-data-calibrated baseline. Without this, 'increased vulnerability to morph-based attacks' in the abstract is broader than the experiment supports.
- [§3.3, Tables 4-7] Tables 4-7, §3.3: extreme quantile estimates are used without uncertainty. At FMR=0.001% (1e-5), Color FERET's 3.7M impostor comparisons imply only ~37 expected false matches, and many observed FMR cells are censored as '<0.0002%'; FNMR=0.000% on FERET is also compatible with zero errors among only 3,652 genuine items. The strong sensitivity claimed at 0.005%/0.001% could be partly estimation noise. Add bootstrap or Wilson/binomial confidence intervals for FNMR/FMR and calibrated quantiles, state effective sample sizes/minimum detectable rates, and mark cells whose uncertainty crosses the target FMR.
- [§3.1, Table 1] Table 1 and §3.1 pair generation: the table is ambiguous about whether 'Gen.' and 'Imp.' are images or pairs. FLUXSyn-ID 44,653 and Syn-Multi-PIE 30,000 look like image counts (approx. IDs × 3 types), whereas the text repeatedly defines genuine pairs as all within-identity comparisons; CFPW is listed with 21,735 'Gen.' despite 500 identities×20 images implying a different pair count. Since impostor pair counts determine quantile resolution and the censored FMR cells, please give exact genuine/impostor pair counts, sampling/exhaustion rules, deduplication, identity-disjointness, and random seeds, and reconcile Table 1 with the protocol.
minor comments (6)
- [Figs. 4-5, §3.3] Figures 4-5 are called 'distance distributions' but the protocol uses cosine similarity; use one term consistently and state explicitly that higher cosine = stricter acceptance threshold.
- [§3.1] §3.1 repeats the sentence defining genuine pairs for each dataset; consolidate once in Pair generation. Also define 'reference', 'illumination_0', 'pose_3', f_doc, live_0e_d1 and live_0p_d1 in a small table rather than inline fragments.
- [§3.2, Table 2] Table 2 mixes LFW accuracy from heterogeneous publications; add a caution that these are not directly comparable and are used only to motivate EdgeFace-S, not to rank models for this study.
- [§3.2-§3.3] Release the threshold-estimation/evaluation scripts, pair lists or hashes, EdgeFace checkpoint identifier, alignment pipeline, and seeds; the protocol is otherwise reproducible but the paper currently reports no artifact.
- [Abstract, §5.2, §7] In §5.2, phrase EES relevance as a proxy argument: FERET/CFPW frontal subsets plus one lightweight model bound controlled/unconstrained extremes but do not include kiosk/mobile sensors, quality filters, demographics, or multi-stage logic; the abstract should mirror this scoping.
- [Throughout] Typos/OCR artifacts in the arXiv text (e.g., 'Fac e', 'verificatio n', ligature breaks) should be cleaned before camera-ready; check that tables' '* <0.0002%' censoring is explained in captions.
Circularity Check
No circularity: standard empirical cross-dataset quantile transfer; results do not reduce to inputs by construction.
full rationale
The paper’s load-bearing chain is purely empirical. Thresholds are defined as quantiles of impostor cosine-similarity scores on a calibration set (§3.3), then applied without image overlap to held-out evaluation sets; FNMR, observed FMR, NMPR, and FMPR are measured directly (Tables 3–8). Nothing is fitted and then re-labeled as a prediction, and no uniqueness theorem or ansatz is imported from the author’s prior work to force the conclusion. Synthetic datasets (FLUXSyn-ID, Syn-Multi-PIE), the EdgeFace model, Color FERET, CFPW, and AMSL morphs are external resources; citing them does not make the transfer failures tautological. The claim that synthetic calibration fails under domain shift because of score-distribution mismatch is an observational finding from those tables, not a restatement of the calibration definition. Correctness concerns (e.g., bulk location/scale shift vs. tail shape, lack of score normalization) are outside the scope of circularity. Score 0 is appropriate.
Assumptions & free parameters
free parameters (3)
- Target FMR operating grid (0.1%, 0.05%, 0.01%, 0.005%, 0.001%) =
0.1% to 0.001%
- Impostor pair sampling size / procedure =
up to ~100M–200M impostor pairs (dataset-dependent)
- EdgeFace-S gamma=0.5 choice =
EdgeFace-S γ=0.5 (3.65M params)
assumptions (4)
- domain assumption Decision thresholds at target FMR equal empirical quantiles of the impostor cosine-similarity distribution on the calibration set.
- domain assumption Cosine similarity on frozen EdgeFace embeddings is an adequate score for studying operational threshold transfer.
- ad hoc to paper Color FERET frontal and CFPW frontal bound controlled vs unconstrained document-to-live conditions relevant to EES-like settings.
- domain assumption Landmark-based AMSL morphs are a meaningful security probe for threshold miscalibration.
Cite this review
Pith. "Pith review of On the Use of Synthetic Data for Threshold Calibration in Face Recognition: Performance and Security Implications for Border Control Systems." pith.science (2026). https://pith.science/paper/WWXCH5OO
@misc{pith2026260725990,
author = {Pith},
title = {Pith review of: On the Use of Synthetic Data for Threshold Calibration in Face Recognition: Performance and Security Implications for Border Control Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/WWXCH5OO}},
note = {Machine review of arXiv:2607.25990}
}
read the original abstract
The recently deployed Entry/Exit System (EES) introduces large-scale biometric verification into European border control, requiring face recognition systems to operate at extremely low false match rates (FMR). While regulatory frameworks define performance targets at the EES Central System level, they do not specify how verification thresholds should be calibrated in practice at the Member State level. In operational settings, obtaining representative real-world data for calibration is often constrained by legal, logistical, and privacy limitations. In this work, we investigate the use of synthetic face data for threshold calibration in document-to-live verification scenarios relevant to border control systems. We analyze the alignment of genuine and impostor score distributions between synthetic and real datasets and evaluate the transferability of calibrated thresholds across domains, with a focus on low-FMR operating points. Our results show that synthetic data can approximate calibration behavior in controlled settings, but fails to reliably generalize to unconstrained conditions due to mismatches in score distribution tails. These discrepancies lead to significant degradation in recognition performance and increased vulnerability to morph-based attacks. We further demonstrate that calibration outcomes are highly dataset-dependent, even across synthetic datasets. Overall, our findings highlight that while synthetic data is useful for system development and preliminary calibration, our results indicate that reliable threshold selection in high-security deployments typically requires validation and adjustment using representative real-world data.
Figures
Reference graph
Works this paper leans on
-
[1]
G. Bae, M. de La Gorce, T. Baltruvaitis, C. Hewitt, D. Chen , J. V alentin, R. Cipolla, and J. Shen. Digiface-1m: 1 million digital face images for face recognition. In IEEE/CVF Win- ter Conference on Applications of Computer Vision (WACV), pages 3526–3535. IEEE, 2023
2023
-
[2]
P . Borsukiewicz, F. Boutros, I. E. Olatunji, C. Beumier, W. C. Ouedraogo, J. Klein, and T. F. Bissyand´ e. Beyond real faces: Synthetic datasets can achieve reliable recognitio n performance without privacy compromise. arXiv preprint arXiv:2510.17372, 2025
arXiv 2025
-
[3]
Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zisserman. Vggface2: A dataset for recognising faces across pose and age. In 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG), pages 67–74. IEEE, 2018
2018
-
[4]
S. Chen, Y . Liu, X. Gao, and Z. Han. Mobilefacenets: Effi- cient cnns for accurate real-time face verification on mobil e devices. In CCBR, 2018
2018
-
[5]
Colbois, T
L. Colbois, T. de Freitas Pereira, and S. Marcel. On the us e of automatically generated synthetic image datasets for benc h- marking face recognition. In IEEE International Joint Con- ference on Biometrics (IJCB) , pages 1–8. IEEE, 2021
2021
-
[6]
DeBruine and B
L. DeBruine and B. Jones. Face research lab london set, 2017
2017
-
[7]
J. Deng, J. Guo, N. Xue, and S. Zafeiriou. Arcface: Ad- ditive angular margin loss for deep face recognition. In 2019 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR) , pages 4685–4694, 2019
2019
-
[8]
Independent eval- uation of biometric technology in the eu: State of play and future options, 2026
eu-LISA, European Commission DG HOME, European Commission JRC, Europol, and Frontex. Independent eval- uation of biometric technology in the eu: State of play and future options, 2026. Policy Brief. Available on- line: https://www.eulisa.europa.eu/our-publications/ policy- brief-independent-evaluation-biometric-technology-eu
2026
Show all 34 references
-
[9]
European Commission. Commission implementing decisio n (eu) 2019/329 laying down the specifications for the quality , resolution and use of fingerprints and facial image for bio- metric verification and identification in the entry/exit sys tem (ees), February 2019. Last accessed: ...
2019
-
[10]
Commission announces launch of shared biometric matching service, May 2025
European Commission. Commission announces launch of shared biometric matching service, May 2025. Last ac- cessed: 23 Apr 2026
2025
-
[11]
Entry/exit system (ees), April 2 026
European Commission. Entry/exit system (ees), April 2 026. Last accessed: 23 Apr 2026
2026
-
[12]
European Parliament and the Council of the European Union. Regulation (eu) 2017/2226 establishing an entry/ex it system (ees) to register entry and exit data and refusal of en - try data of third-country nationals crossing the external b or- ders of the member states, November...
2017
-
[13]
Best practice technical guidelines for automated border contro l (abc) systems, August 2015
Frontex - European Border and Coast Guard Agency. Best practice technical guidelines for automated border contro l (abc) systems, August 2015. Last accessed: 23 Apr 2026
2015
-
[14]
Tech- nical guide for border checks on entry/exit system (ees) re- lated equipment, June 2021
Frontex - European Border and Coast Guard Agency. Tech- nical guide for border checks on entry/exit system (ees) re- lated equipment, June 2021. Last accessed: 23 Apr 2026
2021
-
[15]
George, C
A. George, C. Ecabert, H. O. Shahreza, K. Kotwal, and S. Marcel. Edgeface: Efficient face recognition model for edge devices. IEEE Transactions on Biometrics, Behavior , and Identity Science, 2024
2024
-
[16]
Granoviter, A
O. Granoviter, A. Gruzdev, V . Loginov, M. Kogan, and O. Zvitia. Face recognition using synthetic face data. arXiv preprint arXiv:2305.10079, 2023
2023 arXiv
-
[17]
G. B. Huang, M. Mattar, T. Berg, and E. Learned-Miller. L a- beled faces in the wild: A database for studying face recog- nition in unconstrained environments. In W orkshop on Faces in ’Real-Life’ Images: Detection, Alignment, and Recogni- tion, 2008
2008
-
[18]
Ismayilov, D
R. Ismayilov, D. Sero, and L. Spreeuwers. Fluxsynid: A framework for identity-controlled synthetic face gen- eration with document and live images. arXiv preprint arXiv:2505.07530, 2025
2025 arXiv
-
[19]
A. K. Jain, A. Ross, and K. Nandakumar. Introduction to Biometrics. Springer, 2011
2011
-
[20]
B. F. Klare and A. K. Jain. Heterogeneous face recogniti on: Matching nir to visible light images. In IJCB, 2010
2010
-
[21]
Liu et al
J. Liu et al. Oneface: One threshold for all. In Computer Vision – ECCV 2022 W orkshops , volume 13801 of Lecture Notes in Computer Science . Springer, 2022
2022
-
[22]
Q. Meng, S. Zhao, Z. Huang, and F. Zhou. Magface: A universal representation for face recognition and quality as- sessment. In 2021 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR) , pages 14220–14229, 2021
2021
-
[23]
Neubert, A
T. Neubert, A. Makrushin, M. Hildebrandt, C. Kraetzer, and J. Dittmann. Extended stirtrace benchmarking of biometric and forensic qualities of morphed face images. IET Biomet- rics, 7(4):325–332, 2018
2018
-
[24]
Nguyen, C
K. Nguyen, C. Fookes, S. Sridharan, and S. Denman. Deep learning for face recognition: A survey. arXiv preprint arXiv:1804.06655, 2018
2018 arXiv
-
[25]
Color feret database
NIST. Color feret database. https://www.nist.gov/itl/products-and-services/color-feret- database
-
[26]
Face recognition technology (feret) program
NIST. Face recognition technology (feret) program. https://www.nist.gov/programs-projects/face-recognition- technology-feret
-
[27]
H. Qiu, B. Y u, D. Gong, Z. Li, W. Liu, and D. Tao. Syn- face: Face recognition with synthetic data. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 10860–10870. IEEE, 2021
2021
-
[28]
Schroff, D
F. Schroff, D. Kalenichenko, and J. Philbin. Facenet: A uni- fied embedding for face recognition and clustering. In 2015 IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 815–823, 2015
2015
-
[29]
Sengupta, J
S. Sengupta, J. C. Cheng, C. D. Castillo, V . M. Patel, R. Chellappa, and D. W. Jacobs. Frontal to profile face veri- fication in the wild. In IEEE Conference on Applications of Computer Vision, February 2016
2016
-
[30]
H. O. Shahreza, C. Ecabert, A. George, A. Unnervik, S. Ma r- cel, N. D. Domenico, G. Borghi, D. Maltoni, F. Boutros, J. V ogel, N. Damer, ˜A. Sanchez-Perez, E. Mas-Candela, J. Calvo-Zaragoza, B. Biesseck, P . Vidal, R. Granada, D. Menotti, I. Deandres-Tame, and J. Fierrez. Sdf...
2024
-
[31]
Y . Sun, D. Liang, X. Wang, and X. Tang. Deepid3: Face recognition with very deep neural networks. ArXiv, abs/1502.00873, 2015
2015 arXiv
-
[32]
Wu and K
H. Wu and K. W. Bowyer. What should be balanced in a ”balanced” face recognition dataset? In Proceedings of the British Machine Vision Conference (BMVC) , page 235. BMV A Press, 2023
2023
-
[33]
Yeung, T
M. Yeung, T. Teramoto, S. Wu, T. Fujiwara, K. Suzuki, and T. Kojima. V ariface: Fair and diverse synthetic dataset gener- ation for face recognition. arXiv preprint arXiv:2412.06235, 2024
2024 arXiv
-
[34]
Z. Zhu, G. Huang, J. Deng, Y . Ye, J. Huang, X. Chen, J. Zhu, T. Yang, J. Lu, D. Du, and J. Zhou. Webface260m: A bench- mark unveiling the power of million-scale deep face recog- nition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 2021. IEEE
2021
Reviewed July 31, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.