{"id":"4f45eeb7-8f3d-45e3-96be-0b25b77326d6","arxiv_id":"2411.08490","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Visible iris verification is generally less accurate for dark irises than blue irises, but the effect is model-dependent and one model shows the reverse.","lead":"This study measures how well three iris recognition models work on blue versus dark irises photographed with different smartphones, reporting that verification errors are generally lower for blue irises. It is a check on whether iris biometrics is fair across eye color and devices.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Blue-vs-dark comparison is confounded by unmatched datasets; device and image-quality variation can fully explain the reported EER/TMR gaps.","rationale":"The reader's weakest assumption identifies exactly the load-bearing issue: the blue-vs-dark comparison assumes the datasets are matched in demographics, image quality, and capture conditions. My stress-test confirms this is the most serious threat to the central claim. The paper provides no evidence that blue and dark iris groups are comparable; the huge within-dark device variation (Table 2) demonstrates that device and quality factors can dominate the observed EER/TMR differences. Additionally, the ResNet-50 results contradict the headline 'generally better on blue' even in the only device-matched comparisons, and the text misquotes its own tables (e.g., calling 0.82% a dark-iris EER when Table 2 shows DI-P1 Open-Iris EER=4.86). These issues do not necessarily disprove a real pigmentation bias—the visible-light literature suggests dark irises are harder to image—but the current evidence is not sufficient to establish the claim with the stated generality. The verdict should remain CONDITIONAL until the authors provide a matched or statistically controlled analysis, correct the internal inconsistencies, and report per-color subject counts and quality metrics. No verdict adjustment is needed beyond the reader's conditional acceptance.","tokens_in":10050,"tokens_out":4800,"duration_ms":41200,"concrete_test":"Re-run the evaluation on the original [16] data restricted to device-matched pairs BI-P1 vs DI-P1 and BI-P2 vs DI-P2. For each subset, report per-color subject counts, image-quality metrics (e.g., Laplacian variance, contrast, illumination), and bootstrap 95% confidence intervals for EER and TMR. If the EER differences between blue and dark are not statistically significant, or if the quality metrics differ significantly between blue and dark subsets, the central claim is unsubstantiated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that recognition systems generally perform better on blue irises than dark irises—rests entirely on comparing two blue-iris datasets (P1, P2) with three dark-iris datasets (P1, P2, P3) sourced from [16]. The paper reports 58 unique subjects in total but gives no per-color subject counts, no image-quality statistics, no demographics, and no confidence intervals. The datasets are not matched: blue has no P3 subset, and within P1 and P2 there is no evidence that blue and dark groups were captured under comparable lighting, from comparable subjects, or with comparable image quality. The paper's own Table 2 shows enormous device effects within dark irises (e.g., Open-Iris TMR@1%: DI-P2 43.10 vs DI-P3 94.94), so device and quality variation can easily swamp any pigment effect. The only unconfounded comparisons are BI-P1 vs DI-P1 and BI-P2 vs DI-P2, and in both, ResNet-50 actually performs better on dark irises (EER 19.18/19.53 for blue vs 17.74/17.23 for dark). Thus the 'generally better on blue' conclusion is not supported by the presented evidence; the observed differences could be entirely due to dataset difficulty rather than iris pigmentation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether visible-light iris verification accuracy and fairness vary between blue and dark irises. It uses datasets from [16] captured by three smartphone devices (P1, P2, P3), applies three recognition systems (Open-Iris, ViT-b, ResNet-50), and reports EER, TMR, and fairness metrics DPD and EoD. The authors conclude that recognition systems generally perform better on blue irises, with lower EER and higher TMR, and that fairness varies by device and model.","tokens_in":10214,"tokens_out":4771,"duration_ms":43028,"significance":"The question of pigmentation bias in visible iris recognition is timely and practically relevant, and the paper's stated goal of comparing traditional and deep-learning models across multiple devices is a useful contribution. The paper asks well-defined research questions, benchmarks three systems, and includes fairness metrics, which are appropriate for demographic-bias analysis. However, the central claim is not uniformly supported by the presented evidence: device-matched comparisons are mixed, dataset comparability is unverified, and no uncertainty quantification is provided. If the claim were restricted to model- and device-dependent differences, the study would still be informative but would not support a general statement that visible iris systems perform better on blue irises.","major_comments":[{"comment":"The conclusion that \"recognition systems generally perform better on blue irises\" is not supported by the device-matched comparisons. For ResNet-50 on P1, EER is 19.18% for blue and 17.74% for dark irises; on P2, EER is 19.53% for blue and 17.23% for dark, so dark irises actually perform better in both matched cases. Only Open-Iris and ViT-b favor blue, and the magnitude varies considerably by device. Since the only comparisons that hold device fixed are BI-P1 vs DI-P1 and BI-P2 vs DI-P2, the data support \"model- and device-dependent differences\" rather than a general blue-iris advantage. The conclusion should be revised accordingly.","section":"Section 4.3, Tables 1-2"},{"comment":"The blue and dark iris subsets are not demonstrated to be comparable. The paper reports 58 unique subjects and five device-specific datasets but gives no per-color subject counts, no image-quality statistics, and no demographic information. Device effects within a single iris color are large: for example, Open-Iris TMR at 1% FMR is 43.10% for DI-P2 versus 94.94% for DI-P3 in Table 2. Without matching on image quality, illumination, and subject characteristics, the observed blue-versus-dark differences could be caused by dataset difficulty or capture conditions rather than by iris pigmentation. The authors should either match the blue and dark subsets on quality/difficulty or analyze the device-confounded comparisons separately.","section":"Section 4.1, Tables 2-4"},{"comment":"The discussion misquotes the tables. It states that \"the Open-Iris model achieved an EER of 0.24% for blue irises captured by a P1, whereas it had an EER of 0.82% for dark irises.\" Table 1 reports 0.82% for BI-P2 (blue), not for dark irises; Table 2 reports dark-iris Open-Iris EERs of 4.86%, 8.00%, and 9.16%. The claim that \"Table 2 highlights that dark irises generally exhibit higher error rates compared to blue irises\" is also contradicted by the ResNet-50 rows in the same table. Please correct the discussion to match the data.","section":"Section 4.5"},{"comment":"The evaluation protocol uses a single gallery/probe split with no cross-validation, bootstrap, or significance testing, and all reported metrics are point estimates. With only 58 subjects and roughly five images each, the differences between blue and dark subsets, and especially the fairness metrics in Tables 3-5, need confidence intervals or significance tests to be interpretable. Without such uncertainty quantification, the reported gaps cannot be distinguished from sampling variation.","section":"Section 4.2"},{"comment":"The DPD and EoD fairness metrics are not defined in the paper, and the threshold at which they are computed is only mentioned as \"0.1\" without specifying whether this is FMR, a score threshold, or a normalized threshold. It is also unclear how the fairness comparisons are built from the same gallery/probe splits used for accuracy, and whether the very large values in Table 5 (e.g., DPD 88.22% and EoD 95.98%) are stable given the small sample. Please provide the metric formulas, the threshold basis, and sensitivity analysis or uncertainty estimates.","section":"Section 4.4"}],"minor_comments":[{"comment":"The illustrative performance numbers in Figure 2 (e.g., EER 0.29, TMR 98.80/98.50) do not match any row in Tables 1 or 2. Please either label the figure as an illustrative example or make the values consistent with the reported results.","section":"Figure 2"},{"comment":"The text describes EoD = 27.90% as \"small\" while the paper elsewhere treats higher EoD as indicating greater bias. Please reconcile these statements or define a threshold for what counts as small.","section":"Section 4.4"},{"comment":"The DET curves are not explicitly linked to the model/device rows in Tables 1 and 2. Please ensure each curve is labelled with the model and dataset it represents so the reader can map figures to tables.","section":"Figures 4-5"},{"comment":"There are several typographical and consistency issues, including \"DpD\" versus \"DPD\", \"baised\" in Figure 2, and inconsistent use of decimal points (e.g., \"0.1\" vs \"0.1%\"). A careful proofreading pass would improve clarity.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses an important fairness question and contains useful benchmarking, but the central claim is overstated and the experimental design does not yet rule out major confounds. The paper is within scope for a computer-vision/biometrics venue, but it is currently closer to a workshop-level empirical study. I recommend major revision rather than rejection because the underlying question is significant and the data could be re-analyzed to support a more limited, device- and model-specific conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dave, this one has a real question hiding behind a shaky experiment. The paper wants to show that visible iris verification performs better on blue than dark irises, and it does report lower EER for blue on two of three models. But the evidence doesn't actually support the generalization, and the dataset design leaves a confound that could explain the whole pattern.\n\nWhat's genuinely new: this is the first study to my knowledge that puts blue-versus-dark iris performance side-by-side across three smartphone devices and three model families (classical Open-Iris, ViT-b, ResNet-50), and it adds a fairness analysis using DPD and EoD across devices and colors. The raw measurements are new, and the cross-device fairness tables are a useful addition.\n\nThe soft spots are load-bearing. The datasets come from an existing collection [16], with blue only on phones P1 and P2 and dark on P1, P2, and P3. The paper reports 58 subjects but no per-color subject counts, no image-quality statistics, no demographics, and no check that the blue and dark subsets are comparable in difficulty. Device effects are enormous (Open-Iris TMR@1% goes from 43.10 on DI-P2 to 94.94 on DI-P3), so device/quality variation can swamp pigment effects. The only clean same-device comparisons are BI-P1 vs DI-P1 and BI-P2 vs DI-P2. On those, Open-Iris and ViT-b do favor blue, but ResNet-50 actually does better on dark (EER 19.18/19.53 vs 17.74/17.23). So 'generally perform better on blue' is not what the tables show. The paper also misquotes its own results: the discussion says Table 1 gives an EER of 0.82% for dark irises, but that number is actually for blue on P2; and it says ViT-b gives 'no matches' at 0.1% FMR for P1, while Table 1 lists a TMR of 42.24%. There are no error bars or significance tests anywhere, and no code or data release.\n\nThe fairness analysis is a useful addition, but the threshold choice (0.1% FMR) is not justified and the magnitudes are not accompanied by uncertainty estimates. The abstract and conclusion treat the bias as established, but the paper would need to either match the datasets properly or restrict its claims to model-specific observations.\n\nWho gets value from this? Researchers working on fairness in biometrics or visible-light iris recognition will find the measurements and fairness tables a starting point, but they should not cite the headline without verifying the tables. My recommendation: send it to peer review with a request for major revision. The question is important enough that a referee round is worth the time, but the conclusions in the current form are not supportable.","headline":"The paper's headline claim is not supported by its own tables; the blue-vs-dark comparison is confounded by unmatched datasets, but the raw measurements and fairness tables are worth a second look.","tokens_in":10834,"tokens_out":4812,"would_cite":false,"duration_ms":38366,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Visible-light iris recognition performs better on blue irises than on dark irises.","keywords":["iris recognition","visible spectrum","iris pigmentation","demographic bias","fairness metrics","equal error rate","deep learning","smartphone biometrics"],"falsifier":"Recompute the blue-versus-dark comparison on the same five datasets while controlling for image quality indicators such as focus, contrast, and visible texture, or matching subjects across color groups by age and skin tone; if the blue advantage shrinks to near zero once those controls are applied, the claim that pigmentation itself drives the bias is not supported. A simpler direct check is to rereport per-color subject and image counts and quality statistics from the source datasets.","tokens_in":9795,"feed_emoji":"👁️","tokens_out":4121,"duration_ms":34399,"temperature":0.7,"pith_summary":"This paper sets out to show that iris pigmentation biases visible-light iris verification: recognition models consistently achieve lower Equal Error Rates and higher True Match Rates on blue irises than on dark irises across multiple smartphone capture devices. The authors compare three recognition systems—a classical image-processing pipeline, a fine-tuned Vision Transformer, and a ResNet-50 with an SVM classifier—on five datasets drawn from 58 subjects, two blue-iris datasets and three dark-iris datasets. They also quantify demographic fairness with Demographic Parity Difference and Equalized Odds Difference, finding larger disparities for dark irises, especially across different phones. If the claim holds, visible-light iris verification carries a systematic accuracy disadvantage for people with dark eyes, and fairness auditing of such systems should be standard practice.","feed_headline":"Iris recognition favors blue eyes over dark eyes","feed_subtitle":"Lower error rates for blue irises show up across three models and five smartphone datasets, with fairness gaps up to 88%.","key_machinery":"The load-bearing machinery is a comparative benchmark built from five smartphone-captured iris datasets (two blue, three dark) taken from a single source, processed into normalized iris strips, and fed to three recognition systems: the classical Open-Iris feature extractor, a fine-tuned ViT-b, and a ResNet-50 with SVM classifier. Performance is scored with Equal Error Rate and True Match Rate at two False Match Rate thresholds, and fairness is scored with Demographic Parity Difference and Equalized Odds Difference. These metrics and the shared dataset source are what allow blue-versus-dark and cross-device comparisons to be made at all.","core_discovery":"The central discovery is that visible-light iris verification is not color-neutral. On blue-iris datasets the Open-Iris pipeline achieves EERs of 0.24% and 0.82%, while on dark-iris datasets the same pipeline's EER rises to 4.86%, 8.00%, and 9.16%, with correspondingly lower True Match Rates. The deep models follow the same direction: ViT-b and ResNet-50 show higher EERs and lower TMRs on dark irises than on blue irises, and their absolute performance is generally poor, so the magnitude of the blue advantage depends on model and device. Fairness metrics (DPD and EoD) show the same pattern, with cross-phone dark-iris comparisons reaching DPD values around 65% and same-phone blue-versus-dark comparisons reaching 81% on one device. The paper concludes that recognition systems generally perform better on blue irises, and that pigmentation, model choice, and capture device jointly determine the size of the bias.","pith_inferences":["If the bias is driven by physical contrast rather than algorithmic prejudice, then hardware-side interventions such as active NIR illumination or multi-spectral capture could reduce the gap more directly than model retraining.","The small sample of 58 subjects and the device-iris-color correlation in the dataset leave open the possibility that part of the measured effect is a device-quality artifact; a matched-pair study on a single device with many subjects per color would isolate the pigmentation component.","One testable extension is to generate synthetic dark-iris training data by degrading texture contrast of blue-iris images, and check whether the deep models' dark-iris EER drops toward the blue level.","The same protocol could be applied to NIR-captured irises; if the blue-dark gap disappears under NIR, that would confirm the visible-light optics explanation."],"forward_implications":["Visible-light iris verification systems deployed on smartphones are predicted to reject or downgrade dark-eyed users more often than blue-eyed users, given the same model and capture device.","Fairness metrics like DPD and EoD should be reported alongside EER and TMR whenever iris verification is evaluated across demographic groups.","Training on diverse iris colors helps generalization, but the paper's results indicate that improvement is model- and device-specific, so a single diverse training set does not remove the bias.","For dark-iris recognition, the classical Open-Iris pipeline outperforms the two deep models tested, suggesting that architecture choice matters more in the dark-iris regime."],"supporting_citations":[{"why":"Supplies the five smartphone-captured visible iris datasets and the 58 subjects on which all blue-versus-dark comparisons are made.","marker":"[16]"},{"why":"Provides the Open-Iris recognition pipeline used as one of the three systems.","marker":"[2]"},{"why":"Provides the ResNet-50 architecture used as the deep-CNN-based recognizer.","marker":"[9]"},{"why":"Provides the Vision Transformer (ViT-b) model architecture that is fine-tuned for iris recognition.","marker":"[6]"},{"why":"Defines the Demographic Parity Difference and Equalized Odds Difference fairness metrics used to quantify bias.","marker":"[11]"}],"fun_headline_variants":["Blue eyes get better iris scans than dark eyes","Iris scanners show bias toward blue eyes","Dark iris? Face higher error rates in iris recognition","Iris verification: Blue beats dark in accuracy","Study: Iris recognition less accurate for dark eyes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The datasets compared as 'blue' and 'dark' are assumed to differ essentially only in iris color, with no per-subject or per-image accounting for quality, lighting, or demographics, so any measured performance gap is attributed to pigmentation rather than to the devices or the people photographed.","fun_headline_variants_meta":{"raw":{"variants":["Blue eyes get better iris scans than dark eyes","Iris scanners show bias toward blue eyes","Dark iris? Face higher error rates in iris recognition","Iris verification: Blue beats dark in accuracy","Study: Iris recognition less accurate for dark eyes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0002,"raw_usage":{"total_tokens":1394,"prompt_tokens":985,"completion_tokens":409,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":338}},"tokens_in":601,"tokens_out":409,"duration_ms":4364,"temperature":1.0,"reasoning_tokens":338,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:31:21.984788+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the blue-versus-dark comparison on the same five datasets while controlling for image quality indicators such as focus, contrast, and visible texture, or matching subjects across color groups by age and skin tone; if the blue advantage shrinks to near zero once those controls are applied, the claim that pigmentation itself drives the bias is not supported. A simpler direct check is to rereport per-color subject and image counts and quality statistics from the source datasets.","supporting_citations":[{"cited_title":"Smartphone based visible iris recognition using dee p sparse ﬁltering","cited_arxiv_id":null,"evidence_quote":"Supplies the five smartphone-captured visible iris datasets and the 58 subjects on which all blue-versus-dark comparisons are made."},{"cited_title":"Iris: Iris recognition inference system of the worldcoin project, 2023","cited_arxiv_id":null,"evidence_quote":"Provides the Open-Iris recognition pipeline used as one of the three systems."},{"cited_title":"Dee p residual learning for image recognition","cited_arxiv_id":null,"evidence_quote":"Provides the ResNet-50 architecture used as the deep-CNN-based recognizer."},{"cited_title":"Fairness index measu res to evaluate bias in biometric recognition","cited_arxiv_id":null,"evidence_quote":"Defines the Demographic Parity Difference and Equalized Odds Difference fairness metrics used to quantify bias."}],"review_version":1}