{"id":"c513a9aa-196c-4e74-8d76-123c79fd58d2","arxiv_id":"2607.13515","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"DriveFace is a 70-subject public benchmark pairing VIS smartphone enrollment with NIR through-glass in-vehicle probes, on which current face-recognition models reach only ~8-12% EER under the hardest tint-and-illumination protocol.","lead":"A new public dataset, DriveFace, pairs smartphone enrollment photos with infrared video of 70 volunteers sitting inside cars, filmed through real and simulated automotive glass. Baseline face-recognition systems lose accuracy sharply through dark tint and at low illumination, showing current models fall short in on-the-move border control.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-session pairing is unspecified: if enrollment uses one reference session for outdoor probes and the other for indoor/simulation, Table 4's protocol ranking is confounded by unequal enrollment–probe time gaps.","rationale":"The stress-test pass focused on the strongest claim. The paper's new artifact is a dataset; its credibility rests on the benchmark protocols being correctly specified. The reader identified the under-specified cross-session pairing. I agree this is the single most load-bearing weakness: it is a specific, verifiable omission, not a matter of taste. The concern is not that the dataset is fraudulent; it is that the published protocol description does not establish that the reported performance differences across protocols are due to the intended physical variables rather than temporal confound. The concrete test is feasible from the released filenames/metadata. Secondary issues (no error bars, tint/identity inconsistencies, aggregate AUC vs per-tint AUCs) are worth fixing but less central. My recommendation remains CONDITIONAL: the paper should be accepted only if the protocol pairing is clarified and, if necessary, Table 4 is recomputed under balanced pairing.","tokens_in":13122,"tokens_out":3329,"duration_ms":35830,"concrete_test":"In the released dataset metadata, extract per-identity enrollment reference session (R1/R2) and probe session for each protocol (outdoor, indoor car, simulation). Count how many test identities are enrolled from each reference session per protocol. Then re-run Table 4's EER/VR evaluation with enrollment reference session balanced across protocols (e.g., enroll half the outdoor probes from R1 and half from R2, same for indoor/simulation). If the EER ordering changes or the simulation gap narrows, the original protocol ranking was confounded. If the pairing is already balanced, report this in the paper and the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's cross-session evaluation claim (Sec. 4.1) requires that enrollment and probe samples come from different sessions for every protocol. Sec. 3.5 labels probe session 1 = outdoor, session 2 = indoor, while reference sessions are smartphone captures ~2 months apart. The text never states which reference session is paired with which probe session. If outdoor enrollment always draws from the temporally closer reference session and indoor/simulation from the farther one, the EER differences in Table 4 (outdoor 2.7%, indoor car 3.1%, simulation 8–12%) could reflect appearance-change interval rather than glass/illumination difficulty. This would invalidate the protocol-difficulty ranking and weaken the claim that DriveFace exposes 'clear performance limitations' specifically due to through-glass and cross-spectral conditions. The dataset itself remains valuable, but the benchmark protocol as documented is not unambiguous.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DriveFace, a publicly released dataset for cross-spectral, through-glass face recognition in vehicular border-control scenarios. It pairs VIS smartphone pre-enrollment videos with NIR in-vehicle probes acquired through automotive windows, under outdoor, indoor-car, and indoor simulated-tint protocols, and adds a presentation-attack subset (DriveFace-PAD). The authors provide standardized FR and PAD protocols, evaluate several public baselines (AdaFace, LVFace, EdgeFace, xEdgeFace for FR; DeepPixBiS, CLIP, DinoV2, EfficientNet, ConvNeXtV2 for PAD), and report that performance degrades most in the simulated-tint protocol and in unseen-attack PAD scenarios. The dataset includes rich metadata (tint level, illumination, pose, speed) and is offered as a benchmark for future research.","tokens_in":13195,"tokens_out":4133,"duration_ms":47207,"significance":"If the protocol issues are resolved, DriveFace fills a genuine gap: existing public NIR-VIS and in-vehicle datasets do not jointly capture external-view, through-glass, cross-spectral acquisition with motion and pose variation. The metadata granularity and the inclusion of a through-glass PAD subset are valuable assets. The paper is transparent about poor unseen-attack PAD results, reports standard metrics, and uses external, publicly available baselines, which supports reproducibility. The main scientific contribution is the benchmark itself and the reference numbers it establishes; the observed protocol-ranking and tint-level trends would be useful evidence for the field if presented with appropriate uncertainty quantification.","major_comments":[{"comment":"The cross-session protocol is under-specified and potentially confounded. Sec. 4.1 states that 'samples from different sessions are used for enrollment and probe sets to ensure cross-session evaluation,' but Sec. 3.5 defines probe 'session 1' as outdoor and 'session 2' as indoor, while the two reference sessions are smartphone captures roughly two months apart. The paper never states which reference session is paired with which probe session for each protocol. If enrollment for outdoor probes always uses the temporally closer reference session and enrollment for indoor/simulation probes uses the farther one, the protocol ranking in Table 4 (outdoor EER 2.69%, indoor car 3.11%, simulation 8.09-11.73%) could reflect unequal enrollment-probe time gaps rather than through-glass or illumination difficulty. Please specify the exact pairing and, if needed, re-run with a balanced design, or expl","section":"Sec. 4.1 and Sec. 3.5"},{"comment":"All FR metrics are point estimates computed on only 28 test identities, with no confidence intervals or per-subject variability. The differences used to support the protocol-difficulty ranking, e.g., outdoor EER 2.69% vs. indoor car 3.11%, and the tint-level AUC differences in Table 5 (98.36 vs. 98.72 for T05 vs. T15) are likely within sampling noise. Please provide bootstrap confidence intervals by subject, or report per-protocol and per-tint distributions. Without this, the central quantitative claims about which conditions are most challenging and about monotonic tint effects are not supported beyond the point estimates.","section":"Tables 4 and 5"},{"comment":"The paper's central claim is that 'state-of-the-art models show clear performance limitations under these realistic conditions.' The FR results in Table 4 show strong performance for outdoor (EER 2.69%) and indoor car (EER 3.11%), with clear degradation only in the simulation protocol. If the cross-session pairing issue in Major Comment 1 is real, the simulation-protocol degradation may be inflated by an enrollment-probe time gap, which would weaken the headline claim. Please clarify or temper the claim, or provide additional evidence that the simulation protocol is intrinsically harder independent of session pairing.","section":"Abstract and Sec. 5"}],"minor_comments":[{"comment":"The Table 6 column definitions are unclear: the 'Attack' column exceeds Print+Mask counts (e.g., grandtest train: 3,519 vs. 9,856), yet replay attacks are stated in Sec. 4.2 to be excluded. Clarify whether replay frames are included in the Attack totals, and define each column precisely.","section":"Table 6"},{"comment":"Using the term 'session' for both temporal capture sessions and probe acquisition conditions (outdoor vs. indoor) is confusing. Consider renaming the probe labels to 'condition' or 'protocol' (e.g., P1/P2) to avoid the ambiguity that underlies Major Comment 1.","section":"Sec. 3.5"},{"comment":"The paper says 'the two sessions were collected approximately two months apart' but does not state whether this holds for all 70 subjects or whether some subjects had different intervals. Also, the description of selecting 'up to 10 samples per subject from each video' should clarify whether these are independent frames or a single tracklet; correlated samples can artificially reduce variance.","section":"Sec. 3.1 and Sec. 4.1"},{"comment":"The ethics statement is brief. Please include details about data anonymization, access controls, and intended data-use agreement. In Sec. 4.3, the xEdgeFace adaptation protocol is described only by reference to [14]; provide the key hyperparameters or state that code will be released to ensure reproducibility.","section":"Sec. 7 and Sec. 4.3"}],"recommendation":"major_revision","confidential_remarks":"This is a useful dataset contribution that is likely acceptable after the protocol-pairing specification and uncertainty analysis are addressed. The main risk is that the reported protocol-difficulty ranking, which is central to the paper's narrative, may be an artifact of undocumented session pairing. I do not see a fundamental flaw in the dataset design, and the authors have been transparent about the poor unseen-attack PAD numbers. The revision should focus on documentation precision and statistical grounding rather than new experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"DriveFace is a genuinely new dataset artifact and worth engaging with. It pairs VIS smartphone enrollment with NIR probes acquired through automotive glass from outside the vehicle, under outdoor, indoor-car, and simulated tint conditions, plus a PAD subset with print/replay/mask attacks. No existing public benchmark combines cross-spectral matching, through-glass degradation, motion, and PAD in one protocol family. The baseline work is solid: they use external pretrained models (AdaFace, LVFace, EdgeFace) plus one adapted model (xEdgeFace), and they report the poor unseen-mask PAD numbers without hiding them. That honesty counts.\n\nThe soft spots are real but fixable. The main one is the cross-session pairing. Section 4.1 says samples from different sessions are used for enrollment and probes, but the paper never states which reference session (smartphone captures ~2 months apart) is paired with which probe session (session 1 = outdoor, session 2 = indoor). If outdoor enrollment always draws from the temporally closer reference session and indoor/simulation from the farther one, the protocol ranking in Table 4 could reflect appearance-change interval rather than acquisition difficulty. That would not destroy the paper's central claim — the simulation protocol still shows EERs around 8–12% and Table 5 shows a sensible monotone tint effect within the controlled subset — but it does undermine the specific 'outdoor easier than simulation' comparison as documented. The authors need to state the pairing explicitly.\n\nSecondary issues: only 28 test identities with no confidence intervals; the tint labels in Section 3.3 (30, 45 VLT) don't match Table 5 (T35, clear); the PAD section mentions four illumination levels while Section 3.3 lists three. None of these are load-bearing, but they add up to a camera-ready revision rather than a finished benchmark report.\n\nWho is this for? Biometrics researchers working on cross-spectral FR, through-glass acquisition, and PAD generalization. A serious referee should see it because the dataset fills a real gap and the baseline results are honest. I would send it to peer review, conditional on the authors releasing the data as claimed and specifying the session protocol. My own verdict is not skeptical — the contribution is real, but the documentation needs tightening.","headline":"A genuinely new public dataset for through-glass NIR face recognition with honest baselines, but the cross-session protocol needs to be spelled out before the headline numbers can be trusted.","tokens_in":13863,"tokens_out":3173,"would_cite":true,"duration_ms":31872,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DriveFace, a new public dataset, pairs visible-spectrum smartphone enrollment with near-infrared in-vehicle probes captured through automotive glass, and its baseline evaluations show current face recognition models degrade markedly under t","keywords":["face recognition","cross-spectral","near-infrared","through-glass","border control","presentation attack detection","dataset","on-the-move"],"falsifier":"Inspect the released metadata to determine the exact mapping from reference session to probe protocol. If the outdoor protocol always enrolls from the reference session temporally closest to the outdoor capture while the simulation protocol enrolls from the farther session, the Table 4 ranking collapses; re-evaluating with balanced or randomized session pairing would then be required to validate the benchmark's cross-session claim.","tokens_in":12879,"feed_emoji":"🚗","tokens_out":6509,"duration_ms":65300,"temperature":0.7,"pith_summary":"The paper introduces DriveFace, a public benchmark for face recognition at vehicular border crossings, where travelers are photographed through car windows while approaching a checkpoint. It pairs clean smartphone enrollment images with near-infrared probe images captured from outside real vehicles and from a controlled window-tint simulation, and it includes a presentation-attack subset. The central claim is that current face recognition models show clear performance limitations under these realistic conditions: verification error rates climb to about 8–12% in the simulated tinted-window protocol, and presentation-attack detectors that work on known attacks fail badly on unseen mask and print attacks. The dataset's value is that it provides a metadata-rich, common testbed for developing methods that handle cross-spectral matching, glass tint, motion, and spoofing together—conditions existing benchmarks do not reproduce.","feed_headline":"Tinted-glass face-ID error rates reach 12% in border benchmark","feed_subtitle":"Public dataset pairs smartphone photos with in-car NIR captures to expose where current models fail.","key_machinery":"The central object is the DriveFace dataset itself: 70 consenting subjects captured over two sessions, with visible-spectrum enrollment videos from two consumer smartphones and near-infrared probes from an infrared sensor under three protocols (real-vehicle outdoor, real-vehicle indoor, and indoor simulated window with five tint levels). The benchmark's machinery is the paired acquisition design and per-acquisition metadata (tint, illumination, pose, speed) that enable per-factor analysis. For face recognition, the evaluation uses a 60/40 identity split with cross-session enrollment/probe separation; for presentation attack detection, it defines grandtest, unseen-print, and unseen-mask proto","core_discovery":"The paper's discovery is that existing face recognition technology, including strong pretrained models, is not yet reliable for on-the-move vehicular border control. Using the new DriveFace dataset, the authors show that verification error rates range from 2.69% EER on outdoor clear-glass captures to 8.09–11.73% EER on indoor simulated tinted-window captures, and that recognition accuracy degrades monotonically as window tint darkens. They also show that presentation attack detection models that work well on known attacks fail on unseen attacks, with error rates as high as 45.8% ACER for unseen masks. These results establish DriveFace as a challenging benchmark that isolates the combined eff","pith_inferences":["The controlled simulation protocol, which varies only tint and illumination, could support a quantitative model linking VLT percentage to expected recognition error; one could test whether error rates scale predictably with transmission.","Outdoor real-vehicle captures achieved lower errors than indoor simulation, suggesting natural lighting and motion may hurt less than controlled low illumination and tint; a direct stationary-versus-moving outdoor comparison could isolate the true cost of motion.","The metadata could support a per-factor error decomposition (pose, tint, illumination, speed), letting border authorities choose acquisition policies—camera angle, illuminator power, or tint regulation—that optimize recognition within operational constraints.","The unseen-attack protocols could be extended to 3D-printed or high-fidelity silicone masks; given the very high errors on masks, such attacks would likely remain a serious vulnerability."],"forward_implications":["DriveFace gives researchers a public testbed where cross-spectral matching, glass tint, motion, and illumination can be studied together, something existing public benchmarks do not offer.","The reported error rates (up to about 11.7% EER on the simulation protocol) quantify how far current recognition is from on-the-move border deployment and set reference numbers for future methods.","The monotone improvement in accuracy as window tint lightens shows tint is a first-order controllable factor, implying that operational tint limits could measurably affect recognition accuracy.","The consistent gains of the cross-spectral-adapted model over its RGB-only backbone indicate explicit domain adaptation is a productive direction for this setting.","PAD results show known attacks are nearly solved (ACER as low as 0.50%) but unseen masks remain open (ACER up to 45.8%), steering anti-spoofing research toward generalization rather than known-attack fitting."],"fun_headline_variants":["Face ID errors reach 12% at tinted border checkpoints","New DriveFace dataset exposes face-recognition weak spots","In-car face ID fails on unseen masks 45.8% of the time","DriveFace: benchmark finds face ID unreliable for moving vehicles","Border face-ID benchmark: tint and motion break current models"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the enrollment and probe images are genuinely cross-session in a way that makes protocol comparisons fair; the paper does not state which of the two reference sessions (captured about two months apart) is paired with each probe session, so if the pairing is imbalanced across protocols, the reported differences could be confounded by enrollment–probe time gaps rather than by glass and tint conditions.","fun_headline_variants_meta":{"raw":{"variants":["Face ID errors reach 12% at tinted border checkpoints","New DriveFace dataset exposes face-recognition weak spots","In-car face ID fails on unseen masks 45.8% of the time","DriveFace: benchmark finds face ID unreliable for moving vehicles","Border face-ID benchmark: tint and motion break current models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0006,"raw_usage":{"total_tokens":2607,"prompt_tokens":680,"completion_tokens":1927,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":424,"completion_tokens_details":{"reasoning_tokens":1840}},"tokens_in":424,"tokens_out":1927,"duration_ms":16402,"temperature":1.0,"reasoning_tokens":1840,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T04:57:56.641037+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the released metadata to determine the exact mapping from reference session to probe protocol. If the outdoor protocol always enrolls from the reference session temporally closest to the outdoor capture while the simulation protocol enrolls from the farther session, the Table 4 ranking collapses; re-evaluating with balanced or randomized session pairing would then be required to validate the benchmark's cross-session claim.","supporting_citations":[],"review_version":1}