{"id":"370c9c3e-4e1f-4e4c-b069-ed276b51817b","arxiv_id":"2505.18538","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Multimodal EOG and eye-tracking LSTM classifies simulated refractive power with 96.2 percent accuracy within a person, but only 8.9 percent across people, near chance.","lead":"This paper tests whether eye movement signals from EOG and video eye tracking can reveal a person's refractive error, using LSTM models trained on 37 people wearing 13 trial lenses. Personalized models reached 96 percent accuracy, but models applied to new people scored near chance, so the approach is not yet a general screening tool.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Within-subject accuracy may reflect time-varying recording drift confounded with lens condition, since the paper never states lens order was randomized or that drift was controlled.","rationale":"The reader's weakest assumption already points to eye-tracker instability and lack of recalibration, and I agree that those are serious. My stress-test sharpens this into a concrete, testable confound: the manuscript does not report whether the order of the 13 diopter conditions was randomized or counterbalanced, so slow temporal drift can be perfectly aligned with refractive power. This is more load-bearing than the trial-lens proxy objection because it threatens even the narrower claim about induced refractive states: if the lenses are ordered and the tracker drifts, the model can achieve 96% accuracy by learning the drift, not the optical state. It also explains the multimodal advantage, since the eye-tracking stream contributes 93 features and the EOG stream only 8, so the fused model has many more ways to encode time-dependent artifacts. The paper deserves credit for openly listing the lack of recalibration and unstable confidence as limitations, and for reporting the honest near-chance cross-subject result. Those admissions do not neutralize the confound, but they make the issue a testable methodological gap rather than an internal inconsistency. I would not move the verdict away from CONDITIONAL: the requested conditions should include releasing the raw data or a timestamped feature table, reporting the exact condition order, and running the time-split and confidence-restricted analyses above. If those checks show no drift effect, the central claim would be substantially stronger. If they show a drift effect, the headline accuracy would need to be reinterpreted as a recording-artifact benchmark rather than evidence for refractive-state-dependent eye movements.","tokens_in":21670,"tokens_out":4347,"duration_ms":42025,"concrete_test":"Re-analyze the subject-dependent data with condition order as a strict confound. First, determine from the raw dataset whether lens order was randomized; if it was fixed, this alone is a red flag. Second, train the same LSTM with a session-time index or condition-order channel added as an input and compare to the reported 96.2% accuracy; if a time-only model reaches comparable accuracy, the refractive-state interpretation fails. Better: train on the first half of each participant's session and test on the second half (and vice versa); if accuracy drops to near chance, the classifier relied on temporal drift rather than lens identity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing hole is not the trial-lens proxy per se but a temporal confound that would explain the headline number even if the lenses had no refractive effect. The manuscript never states that the order of the 13 lens conditions was randomized or counterbalanced across participants; Section 3.1 only randomizes target positions and pursuit trajectories within a condition. The eye tracker was calibrated once at session start and never recalibrated (Section 7), and the authors concede that small positional offsets may have accumulated and that discriminative patterns may reflect fluctuations in device stability (Sections 5-6). Under a fixed or partially ordered condition sequence, any slow drift in gaze offset, pupil baseline, EOG electrode polarization, or participant state is perfectly correlated with diopter. A four-layer LSTM with 512 hidden units can readily memorize such drift, yielding near-ceiling within-subject accuracy. This also explains why the feature-rich multimodal model (93 eye-tracking + 8 EOG features) beats both unimodal models: more features provide more drift signatures. Thus the central claim that the signals contain refractive-state-dependent information is not established until the temporal-order/drift confound is removed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains unidirectional multi-layer LSTM classifiers to estimate refractive power from EOG and video-based eye-tracking signals in 37 participants whose refractive state was manipulated with trial lenses from -3.0D to +3.0D in 0.5D steps (13 classes). Models are evaluated in subject-dependent (per-user cross-validation) and subject-independent (leave-one-subject-out) settings. The multimodal model reaches 96.207% mean accuracy in the subject-dependent setting and significantly outperforms unimodal EOG (84.451%) and eye-tracking (92.432%) models on a Friedman test with Wilcoxon post-hoc comparisons. In the subject-independent setting, all models perform near the 7.692% chance level (multimodal 8.882%), with no significant between-model differences. The authors interpret the results as evidence for the potential and limitations of passive refractive power estimation from eye movement data.","tokens_in":21903,"tokens_out":2552,"duration_ms":23708,"significance":"If the subject-dependent result reflects genuine refractive-state-dependent changes in eye movements, the paper would make a useful contribution to passive vision screening, particularly by showing that combined EOG and eye tracking can be more informative than either modality alone. The manuscript is careful in several respects: the protocol is described in enough detail to be reproduced, the evaluation covers both personalized and cross-subject scenarios, statistical testing is appropriate for the comparisons made, and the limitations (especially eye-tracker non-recalibration and low gaze confidence for some participants) are explicitly acknowledged. However, the central quantitative claim is currently threatened by a temporal confounding that the manuscript itself identifies but does not resolve; therefore the significance of the result cannot be assessed until this is addressed.","major_comments":[{"comment":"The manuscript never states whether the order of the 13 lens conditions was randomized or counterbalanced across participants. Because the eye tracker was calibrated once at session start and not recalibrated (Section 7), and because EOG electrodes and participant state can drift over time, any slow temporal drift in signal quality is perfectly correlated with diopter under a fixed or partially ordered condition sequence. A four-layer LSTM with 512 hidden units can readily memorize such drift, which would explain the near-ceiling 96.207% subject-dependent accuracy even if the lenses had no refractive effect. This is load-bearing for the paper's central claim, so the authors should report the condition ordering and provide control analyses, such as label-shuffled baselines, regression on session time, comparison of early versus late session conditions, or a calibration-offset monitoring analysis.","section":"§3.1, §7"},{"comment":"The acknowledged gaze confidence instability for several participants (P21: 0.29 ± 0.39, P22: 0.30 ± 0.35) and the authors' concession that discriminative patterns may reflect fluctuations in device stability directly undercut the interpretation of the eye-tracking and multimodal results. Confidence values were excluded from the features, but the underlying instability can still propagate into pupil and gaze feature values. To support the claim that the models capture refractive-state-dependent signals rather than recording artifacts, the authors should show that the main accuracy results remain when restricted to participants with high and stable gaze confidence (e.g., P5, P30), or after regressing out time and confidence-related effects.","section":"§5, §6, Table 6"},{"comment":"The abstract and Section 4 describe the subject-independent accuracy (8.882%) as 'marginally above chance,' but no statistical test against the chance level of 7.692% is reported. The Friedman test in Table 5 only compares the three models and yields p = 0.482. Since the subject-independent result is used to characterize generalization, the authors should either add an appropriate test against chance (e.g., binomial or permutation test across subjects) or temper the claim to a descriptive observation.","section":"Abstract, §4, Table 5"}],"minor_comments":[{"comment":"The text has several capitalization inconsistencies (e.g., 'We employ' mid-sentence) and the sentence beginning 'To fill this gap, We employ' should be revised.","section":"§3.3"},{"comment":"The dataset is described as publicly available, but reference [52] is the authors' own work; please clarify the dataset's public availability and provide a URL or repository link if applicable.","section":"§3.1, Reference [52]"},{"comment":"Figure 3 reports error bars as standard error of the mean while Table 2 reports ± values without specifying whether they are standard deviations or standard errors; please standardize the notation and caption descriptions.","section":"Figure 3, Table 2"},{"comment":"The statement that participants with high-quality eye-tracking data also achieved strong classification performance could be quantified (e.g., correlation between mean gaze confidence and accuracy) to support the argument.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's main result is interesting but currently vulnerable to a temporal-drift confound that the authors themselves acknowledge. The revision should focus on demonstrating either that lens order was randomized/counterbalanced or that the reported accuracies are robust to time-related drift. If the authors cannot supply such evidence, the central claim should be substantially weakened, and the paper might be better framed as a feasibility study with explicit caveats about artifact-driven performance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely new application—classifying induced refractive power from EOG plus video eye tracking—but the headline within-subject accuracy is not yet trustworthy because of a temporal confound the paper itself half-discloses. The cross-subject result (8.9%, chance 7.7%) is honestly presented as a failure, and the statistical reporting is above average for this kind of paper.\n\nWhat is new: no cited prior work combines EOG and video eye tracking for refractive power classification, and the eye-tracking-only passive classification is new. The experimental protocol is described in enough detail to see what was done: 37 participants, 13 trial-lens conditions, 8-fold trial-segment CV for within-subject, leave-one-subject-out for cross-subject. The Friedman/Wilcoxon analysis is appropriate, and the authors print per-subject accuracies and gaze confidence rather than hiding bad subjects.\n\nThe soft spots are serious, and they all land on the central positive claim. The paper never states that the order of the 13 lens conditions was randomized or counterbalanced. The eye tracker was calibrated once at session start and not recalibrated afterward (Section 7). If the lens sequence was fixed or ordered, any slow drift in gaze offset, electrode polarization, or participant state is collinear with diopter, and a 4-layer LSTM with 512 hidden units has plenty of capacity to memorize that drift. The authors themselves note that discriminative patterns may reflect fluctuations in device stability rather than true differences in eye movement, and admit low, unstable confidence for P21 and P22. So the within-subject 96% may be an artifact of recording drift, not refractive state. This is not a minor caveat; it is the central issue.\n\nSecond, the paper claims multimodal fusion still highlights the advantage of data fusion in the subject-independent setting, but the Friedman test across the three models gives p = 0.482. The data do not support that sentence. Third, the modality comparison is not clean: 93 eye-tracking features versus 8 EOG features, so the multimodal gain over EOG could just be feature count. The authors acknowledge this imbalance. Fourth, the dataset is called publicly available but is the authors' own [52], and no code or data link appears; reproducibility is not yet possible.\n\nWho should read it: people working on passive eye-movement monitoring and on confound control in physiological ML. The negative cross-subject result is useful calibration. But the central claim—that eye movement signals carry refractive-state information—is not established until the lens-order/drift confound is addressed, the tracker is recalibrated per condition, and the analysis is redone on real refractive error or at least with drift controls and sensitivity checks.\n\nRecommendation: send to peer review, but conditional on major revision that removes the confound or shows it does not explain the result. As it stands, do not cite the 96% as evidence.","headline":"A transparent feasibility study whose within-subject 96% accuracy is likely inflated by eye-tracker drift and unstated lens ordering; the cross-subject failure is honestly reported.","tokens_in":22461,"tokens_out":4098,"would_cite":false,"duration_ms":35029,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that combining EOG and eye-tracking signals lets an LSTM classify an individual's induced refractive power with 96.207% mean accuracy, while cross-person accuracy barely clears chance.","keywords":["refractive error estimation","electrooculography","eye tracking","multimodal learning","LSTM classification","diopter classification","passive vision screening","subject-dependent vs subject-independent generalization"],"falsifier":"Rerun the same LSTM protocol with the eye tracker recalibrated immediately before every lens condition and with chinrest position verified; if the multimodal subject-dependent accuracy falls substantially below 96%, the original result was inflated by condition-correlated device offsets. A second decisive check is to compare accuracy on participants with high versus low gaze confidence: if the model is accurate only for high-confidence participants, low-confidence recordings are not carrying the claimed refractive-state signal.","tokens_in":21476,"feed_emoji":"👁","tokens_out":5521,"duration_ms":45250,"temperature":0.7,"pith_summary":"This paper asks whether refractive power can be estimated passively from eye movements, without the user answering questions or pressing buttons. Using a public dataset of 37 participants who wore trial lenses spanning -3.0 to +3.0 diopters, the authors trained LSTM classifiers on electrooculography (EOG), video eye tracking, and their combination. In a per-person setting, the combined model reached 96.207% mean accuracy across 13 lens conditions, significantly above both single-modality models. In a cross-person setting, all models fell to about chance, showing that the learned patterns are highly person-specific. The paper's contribution is evidence that eye movement signals carry usable information about induced blur, together with a clear picture of where that information does not yet generalize.","feed_headline":"Eye signals reveal lens power 96% of the time per person","feed_subtitle":"Per-user LSTM reads induced refractive power from EOG and gaze, but cross-person accuracy stays near chance.","key_machinery":"The machinery is a four-layer unidirectional LSTM with 512 hidden units per layer, fed by an input projection with ReLU and a 0.5 dropout, trained for 250 epochs with cross-entropy loss. Inputs are 101 time-aligned features: four EOG channels plus their sample-to-sample differences (downsampled from 512 Hz to 120 Hz), and 93 eye-tracking features covering pupil geometry, gaze direction and position, fixations, and blink type, with low-confidence samples replaced by NaN before Hampel and median filtering. Temporal alignment across the two recording systems is achieved with event triggers marking trial start and end. In the subject-dependent setting, trial segments from one participant are cross-validated; in the subject-independent setting, leave-one-subject-out is used. The LSTM's role is to capture the temporal dynamics of eye movement that differ across lens conditions, and the multimodal input is what the paper argues supplies complementary information: EOG's fine-grained, stable voltage signals and the eye tracker's richer gaze and pupil features.","core_discovery":"The central claim is that fusing EOG with video-based eye-tracking features yields a personalized classifier that can read an individual's induced refractive state from how they move their eyes. In the subject-dependent scenario, the multimodal LSTM achieved 96.207% mean classification accuracy over 13 diopter conditions, outperforming eye tracking alone (92.432%) and EOG alone (84.451%), with the differences statistically significant (Friedman chi2 = 55.62, p < .001; pairwise Wilcoxon p < .001). The confusion matrices show that residual errors are not random: the model most often confuses lenses of the same magnitude with opposite sign, such as -2.0 D and +2.0 D. In the subject-independent scenario, the same architecture produced only 8.882% mean accuracy, statistically indistinguishable from the 7.692% chance level, and the authors attribute the gap to known difficulties of cross-person physiological signals. They conclude that eye movement data support personalized, passive refractive-power monitoring but do not yet support a generalizable screening model.","pith_inferences":["If the eye tracker had been recalibrated before each lens condition, the subject-dependent eye-tracking accuracy might drop, which would indicate that part of the signal was device drift rather than refractive state; this is a direct test of the paper's core interpretation.","Collapsing the 13 classes into coarser groups, such as blur magnitude or a simple 'blurred vs. clear' distinction, might survive cross-person transfer better, given that even the per-person errors concentrate on sign.","Because EOG alone already reaches 84% per-person accuracy with only eight features, a minimal wearable EOG setup might be sufficient for longitudinal personal monitoring, with eye tracking adding marginal value only when its calibration is reliable.","The very low gaze confidence of some participants (e.g., P21 at 0.29 +/- 0.39) raises the possibility that per-person accuracy is carried by the high-confidence majority; checking accuracy stratified by gaze confidence would separate true signal from tracking-quality artifacts."],"forward_implications":["A personalized eye movement model could monitor refractive state passively over time, since within-subject classification accuracy is high across all 13 lens conditions.","Eye tracking alone is a stronger single modality than EOG for this task in the per-person setting, suggesting gaze and pupil features carry more discriminative blur-related information.","The systematic confusion between positive and negative lenses of equal magnitude implies that models may first learn blur strength and only secondarily blur sign.","Any practical deployment for screening would need per-user calibration or some form of domain adaptation, because untouched cross-person accuracy is essentially chance.","The multimodal advantage, although modest in the cross-person setting, persists there (8.882% vs. 8.640% and 7.936%), which the paper reads as evidence that fusion remains the more promising foundation."],"supporting_citations":[{"why":"Supplies the dataset of EOG and eye-tracking recordings under 13 trial-lens conditions that all models are trained and evaluated on.","marker":"[52]"},{"why":"Provides the prior EOG-only refractive-power classification approach whose subject-dependent and subject-independent pattern this study extends to eye tracking and fusion.","marker":"[51]"},{"why":"Establishes the eye-tracking route to refractive error estimation via continuous psychophysics, the active-task baseline this passive method contrasts with.","marker":"[40]"},{"why":"Shows EOG signal patterns can classify refractive-error groups with data mining, the EOG baseline that motivates the current approach.","marker":"[22]"},{"why":"Documents the eye tracker whose pupil, gaze, blink, and fixation streams form the eye-tracking feature set.","marker":"[21]"},{"why":"Supplies the LSTM architecture adapted here for temporal classification of physiological signals.","marker":"[38]"}],"fun_headline_variants":["Per-person eye motion reads lens power with 96% accuracy","Fusing EOG and gaze nails lens power per person, but not across people","Eye-movement model passes personalized lens test, fails general screening","Passive eye tracking predicts refractive power, but only for the individual","Per-person 96% accuracy, cross-person near chance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that trial lenses worn by healthy participants reproduce the eye movement effects of genuine refractive error, and that the high within-person accuracy reflects those refractive-state-dependent changes rather than recording artifacts such as the eye tracker's unrecalibrated head-position offsets and unstable gaze confidence.","fun_headline_variants_meta":{"raw":{"variants":["Per-person eye motion reads lens power with 96% accuracy","Fusing EOG and gaze nails lens power per person, but not across people","Eye-movement model passes personalized lens test, fails general screening","Passive eye tracking predicts refractive power, but only for the individual","Per-person 96% accuracy, cross-person near chance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000942,"raw_usage":{"total_tokens":4050,"prompt_tokens":998,"completion_tokens":3052,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":2962}},"tokens_in":614,"tokens_out":3052,"duration_ms":20046,"temperature":1.0,"reasoning_tokens":2962,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:28:52.331217+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the same LSTM protocol with the eye tracker recalibrated immediately before every lens condition and with chinrest position verified; if the multimodal subject-dependent accuracy falls substantially below 96%, the original result was inflated by condition-correlated device offsets. A second decisive check is to compare accuracy on participants with high versus low gaze confidence: if the model is accurate only for high-confidence participants, low-confidence recordings are not carrying the claimed refractive-state signal.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the dataset of EOG and eye-tracking recordings under 13 trial-lens conditions that all models are trained and evaluated on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior EOG-only refractive-power classification approach whose subject-dependent and subject-independent pattern this study extends to eye tracking and fusion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the eye-tracking route to refractive error estimation via continuous psychophysics, the active-task baseline this passive method contrasts with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows EOG signal patterns can classify refractive-error groups with data mining, the EOG baseline that motivates the current approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the eye tracker whose pupil, gaze, blink, and fixation streams form the eye-tracking feature set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the LSTM architecture adapted here for temporal classification of physiological signals."}],"review_version":1}