{"id":"faff9611-824a-4bde-960c-3c75b8c7e1e7","arxiv_id":"2607.09883","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Monte Carlo simulation of a joint gaze-plus-pupil synchronization metric shows theoretical separability of genuine ocular responses from replay, deepfake, and mechanical spoofs when rendering latency is non-zero.","lead":"A dual-stream challenge that randomizes both target motion and brightness is proposed to catch ocular spoofs by checking whether gaze and pupil responses stay locked to biological latencies. Monte Carlo runs on literature latency numbers show the joint score separates genuine eyes from replay, deepfake, and mechanical attacks when rendering lag is non-zero.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the paper's own caveats; the reader's weakest assumption already captures the load-bearing limit.","rationale":"The strongest claim is carefully scoped as simulation-supported theoretical separability, not empirical security. The reader's weakest assumption (non-zero, effectively i.i.d. \tau_render) is exactly the floor the paper quantifies and caveats. The additional idealizations (Gaussian \rho proxy, provisional weights, literature means without subject variance) are real but secondary and already labeled future work; they do not overturn the internal logic or the multi-round trade-off under the stated generative model. Therefore the CONDITIONAL verdict with high confidence remains appropriate; no adjustment is warranted.","tokens_in":8372,"tokens_out":411,"duration_ms":5055,"concrete_test":"Re-run the N=10 000 Monte Carlo of Section 5.1 replacing the Gaussian-decay approximation for \rho with numerical integration of Eq. 4 (plus Eq. 3) under the same \tau_render sweep; if AUC at 25–50 ms shifts by more than ~0.05 or the multi-round ordering in Table 1 reverses, the reported separability depends on the approximation rather than the latency model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is theoretical separability under Monte Carlo with literature latencies (Section 5, Table 1). The paper itself states that deepfake AUC collapses to chance when τ_render \to 0 and that multi-round gains require i.i.d. lag (Sections 5.2, 6.2). Correlation terms are approximated by Gaussian decay of latency deviation rather than full integration of the PLR ODE (Eq. 4) or measured pursuit noise; weights and \theta are provisional. These are modeling idealizations the authors flag, not hidden inconsistencies. No stronger internal flaw is required for the claim as written.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes a Spatio-Luminance Sensor Fusion protocol for ocular presentation-attack detection: a dual-stream challenge that simultaneously randomizes a target’s spatial trajectory and luminance, then scores the joint temporal coupling of smooth-pursuit gaze and pupillary light reflex via a Joint Synchronization Metric S_joint (Eq. 5) that combines two max-correlation terms with a biological-latency penalty Φ (Eq. 6). Monte Carlo simulation (N = 10 000 trials per condition) grounded in literature means μ_g = 150 ms and μ_p = 290 ms shows strong separability for video-replay and mechanical/prosthetic attacks (AUC ≈ 0.996–1.0) and latency-dependent separability for generative deepfakes (AUC rising from chance at τ_render = 0 to ≥ 0.997 beyond 75 ms). A multi-round extension (Table 1) further improves deepfake detection when a non-zero, i.i.d. rendering lag is present. The authors explicitly frame the work as a simulation-supported theoretical framework and list human-subject calibration, hardware benchmarking, and adversarial red-teaming as required future work.","tokens_in":8611,"tokens_out":1008,"duration_ms":9333,"significance":"If the theoretical separability holds under real ocular dynamics, the dual-stream challenge supplies a concrete, latency-coupled alternative to single-cue PLR or gaze PAD and directly addresses the growing threat of real-time generative deepfakes. The explicit threat model (replay, deepfake, mechanical), the quantified multi-round trade-off (Table 1), and the transparent statement of the τ_render floor are useful contributions for the PAD community. Strengths include literature-grounded latency priors, an open acknowledgment that the framework collapses when rendering lag approaches zero, and a clear roadmap for empirical validation. The result remains provisional until human-subject and real-pipeline experiments are performed, but the in-silico analysis is a legitimate first step for a systems-security paper.","major_comments":[{"comment":"§5.1: Correlation terms inside S_joint are replaced by a Gaussian decay of latency deviation rather than being computed from simulated time series that integrate the PLR ODE (Eq. 4) and the pursuit model (Eq. 3). While the approximation is internally consistent and produces the claimed AUCs, it leaves open whether realistic hippus, pursuit noise, and first-order recovery dynamics would preserve the reported separability; a short ablation that scores full synthetic trajectories would strengthen the central claim.","section":null},{"comment":"§4.3 and §5: Weights w1 = w2 = 1, w3 = 2 and the decision threshold θ are provisional and never calibrated to a target FAR. Because S_joint is the sole decision statistic, the absolute AUC numbers (and the multi-round gains in Table 1) remain sensitive to these free parameters; the manuscript should either fix them by a stated operating-point criterion or demonstrate that the ranking of genuine vs. attack scores is robust across a plausible weight range.","section":null}],"minor_comments":[{"comment":"Fig. 2 caption and §5.2: the phrase “reliable-detection threshold (AUC ≥ 0.90)” is used without defining the operating point; clarify whether it is the 95th-percentile genuine threshold mentioned earlier or a pure AUC cut-off.","section":null},{"comment":"Eq. (5): the predicted pupil diameter ˆD(t|Ls) is obtained by “integrating the PLR equation,” yet the Monte Carlo never performs that integration; a one-sentence note reconciling the definition with the Gaussian-decay approximation would avoid reader confusion.","section":null},{"comment":"§2: the claim that “no current methodology fuses continuous spatial tracking with simultaneous pupillary luminance response” is strong; a brief check against recent commercial or patent literature on joint pupil–corneal models would make the novelty statement more precise.","section":null},{"comment":"Table 1: report standard errors or confidence intervals on the AUC estimates so that the improvement from M = 1 to M = 10 can be assessed statistically.","section":null},{"comment":"Notation: Σ_g is introduced in Eq. (3) but later only σ_g appears; make the isotropic-variance assumption explicit.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a clean theoretical/systems proposal whose own caveats already bound its claims. It is appropriate for a venue that accepts simulation-first PAD frameworks provided the authors tighten the two modeling idealizations noted above. No citation or novelty-disclosure concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a simulation-only dual-stream PAD idea for iris/ocular systems: randomize both target path and luminance at once, then score joint temporal lock of smooth pursuit and PLR via S_joint. The multi-round AUC-vs-τ_render table is the useful new bit.\n\nWhat is actually new is the simultaneous challenge design and the Synchronization Matrix that couples the two streams with an explicit biological-latency penalty. Prior work treated PLR and gaze separately, or did spatio-temporal fusion for face/gait; this paper correctly flags that gap and fills it with a clean state-space formulation. The Monte Carlo (N=10k, literature means μ_g=150 ms / μ_p=290 ms, explicit replay/deepfake/mechanical models) is internally consistent and produces the claimed curves. Table 1 is the practical takeaway: multi-round averaging helps only when a non-zero rendering lag exists, and the authors state the floor plainly when τ_render→0. Citations are appropriate and not padded. No circular fitting of the metric to the AUCs.\n\nSoft spots are exactly the ones the paper owns. Correlation terms are Gaussian-decay approximations of latency deviation rather than full time-series integration of the PLR ODE or measured pursuit noise. Weights, σs, and θ are provisional. Everything is idealized attack models, not real generative pipelines or human subjects. That does not break the theoretical claim; it just means the security numbers are existence proofs under stated assumptions, not field performance. The human-proxy relay limitation is also stated cleanly.\n\nThis is for people already working on biometric PAD or challenge-response liveness who want a concrete dual-stream design to build on or red-team. It is not a systems paper and not a dataset paper. I would send it to peer review: the framing is honest, the math is transparent, and the multi-round latency trade-off is worth having in the literature even if referees demand the human calibration the authors already list as future work. Engage if you care about ocular PAD; skip if you only want empirical attack results.","headline":"Clean simulation paper that jointly randomizes gaze trajectory and luminance for ocular PAD; results are theoretical separability under literature latencies, not a deployed detector.","tokens_in":9162,"tokens_out":511,"would_cite":false,"duration_ms":5652,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A dual-stream visual challenge can separate live eyes from deepfakes and replays by locking gaze and pupil responses to randomized space and light.","keywords":["Presentation Attack Detection","Ocular Biometrics","Liveness Detection","Sensor Fusion","Deepfakes","Pupillary Light Reflex","Smooth Pursuit","Challenge-Response"],"falsifier":"Run the multi-round protocol against a real-time generative deepfake pipeline whose measured end-to-end rendering latency is driven below roughly 25 ms (or is strongly correlated across rounds) and check whether S_joint still separates from genuine human responses at AUC ≥ 0.90.","tokens_in":9248,"feed_emoji":"👁️","tokens_out":933,"duration_ms":8823,"temperature":0.7,"pith_summary":"Ocular biometrics such as iris recognition are under attack from video replays and real-time generative deepfakes that defeat static liveness checks. Existing defenses usually test only one physiological cue at a time—smooth-pursuit gaze tracking or the pupillary light reflex—so each can be spoofed independently. This paper proposes a Spatio-Luminance Sensor Fusion protocol that issues a single randomized challenge whose target both moves unpredictably and flickers in brightness, then scores whether the eye’s gaze lag and pupil constriction lag stay synchronized with the expected biological latencies. Monte Carlo simulations using literature-derived latency distributions show that genuine responses separate cleanly from replay, mechanical, and deepfake conditions once the attacker’s rendering lag is non-zero, and that repeating the challenge in multiple rounds further improves deepfake detection. The work is offered as a theoretical, simulation-supported framework whose practical thresholds still require human-subject calibration.","feed_headline":"Dual-stream eye challenge spots deepfakes via lag","feed_subtitle":"Randomized motion plus light forces gaze and pupil to stay locked; multi-round trials raise detection once rendering lag exists.","key_machinery":"The Synchronization Matrix and its Joint Synchronization Metric S_joint: a weighted sum of sliding-window Pearson correlations for gaze and pupil streams, minus a quadratic biological-plausibility penalty Φ on the two estimated latencies (μ_g ≈ 150 ms, μ_p ≈ 290 ms).","core_discovery":"A simultaneous dual-stream challenge that randomizes both spatial trajectory and luminance intensity, scored by a Joint Synchronization Metric S_joint that cross-correlates observed gaze and pupil responses against expected biological latencies, yields theoretical separability of live eyes from video-replay, generative-deepfake, and mechanical spoofs under literature-derived latency models, with multi-round averaging raising deepfake detection when a non-zero rendering lag exists.","pith_inferences":["If generative rendering latency continues to fall, pure timing analysis will eventually fail and the protocol will need complementary texture or 3-D shape checks that the paper already flags as out of scope.","The same dual-stream idea could be ported to other biometric channels that possess both a voluntary tracking response and an autonomic reflex (e.g., head pose plus heart-rate photoplethysmography).","Hardware refresh-rate and IR-frame-rate requirements implied by the latency windows become a concrete engineering specification for next-generation iris sensors."],"forward_implications":["Isolated PLR or gaze liveness checks can be replaced or augmented by a single coupled challenge that forces both streams to stay phase-locked to the live stimulus.","Multi-round challenge designs can push the reliable-detection threshold for deepfakes down from ~50 ms to ~25 ms rendering lag, at the cost of longer sessions.","Video-replay and rigid mechanical-aperture attacks become near-perfectly separable once the joint metric is used, because they either decorrelate completely or mismatch the biological pupil waveform.","The framework is intended as an additional timing layer on top of existing iris or face pipelines rather than a standalone biometric."],"fun_headline_variants":["Dual-stream challenge links gaze and pupil to catch ocular deepfakes","Randomized path and light force eye responses to reveal spoof lag","Sync matrix of pursuit and PLR separates live eyes from deepfakes","Multi-round dual challenges boost deepfake detection via rendering delay","Simultaneous gaze-pupil challenge theoretically spots video and generative spoofs"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Deepfake detection works only if the attacker’s pipeline still incurs a non-zero rendering lag that shifts the observed latencies outside the normal biological window; near-zero or highly correlated lag across rounds collapses the score to chance.","fun_headline_variants_meta":{"raw":{"variants":["Dual-stream challenge links gaze and pupil to catch ocular deepfakes","Randomized path and light force eye responses to reveal spoof lag","Sync matrix of pursuit and PLR separates live eyes from deepfakes","Multi-round dual challenges boost deepfake detection via rendering delay","Simultaneous gaze-pupil challenge theoretically spots video and generative spoofs"]},"model":"grok-4.5","effort":"low","cost_usd":0.003834,"raw_usage":{"total_tokens":1237,"prompt_tokens":801,"num_sources_used":0,"completion_tokens":95,"cost_in_usd_ticks":38340000,"prompt_tokens_details":{"text_tokens":801,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":341,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":801,"tokens_out":95,"duration_ms":3051,"temperature":1.0,"reasoning_tokens":341,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T14:50:17.852196+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the multi-round protocol against a real-time generative deepfake pipeline whose measured end-to-end rendering latency is driven below roughly 25 ms (or is strongly correlated across rounds) and check whether S_joint still separates from genuine human responses at AUC ≥ 0.90.","supporting_citations":[],"review_version":1}