{"id":"5ad0e3ae-7cd9-4c2a-9ff7-33bda51cfbb8","arxiv_id":"2501.00597","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Gaze prediction accuracy varies strongly across subjects, and per-subject fixation noise and saccade velocity correlate with prediction error across three different model architectures.","lead":"This paper measures how gaze prediction error varies across individual people for three different prediction models, and finds that noisy fixations and fast saccades are associated with larger prediction errors. The findings suggest that gaze prediction systems, such as those used in foveated rendering, may need to account for per-subject oculomotor differences rather than relying only on average performance.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Saccade-velocity finding may be confounded by saccade amplitude; Table 2 lacks amplitude control.","rationale":"The reader's conditional verdict is reasonable. I identify one additional load-bearing concern not emphasized in the reader's weakest_assumption: the saccade-velocity correlation in Table 2 may be confounded by saccade amplitude. The features PkVelDurRatioRMd and MnVelRMd both scale with saccade size via the main sequence; Fig. 2a shows all models have larger errors for larger saccades. Without partialling out per-subject amplitude or matching amplitude across subjects, the claim that a per-subject velocity trait predicts prediction difficulty is not established. This is particularly important because the fixation-noise half of the claim is acknowledged by the authors as expected ('obviously, the more fixation noise, the harder it is to predict'), so the saccade result is the non-tautological load-bearing component. The paper also lacks per-subject sample counts and CIs (Section 4.2), and the code/data link in Section 3.4 is a placeholder, but these are secondary to the confound. The suggested partial-correlation and amplitude-band analyses would settle the matter. If the correlation survives amplitude control, the central claim stands; if not, the saccade conclusion should be reframed as an amplitude effect.","tokens_in":10936,"tokens_out":6821,"duration_ms":67435,"concrete_test":"Recompute the two saccade rows of Table 2 after adding each subject's median large-saccade amplitude as a covariate (Spearman partial correlation), and also recompute the correlation using only saccades in a narrow amplitude band, e.g., 10-20 dva as in Fig. 2b. If the partial rs drops to near zero or the within-band correlation is not significant, the saccade-velocity claim is confounded by saccade size. Report the simple Spearman correlation between each velocity feature and median amplitude as part of the check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 and Table 2 report that subjects with higher saccade velocities (PkVelDurRatioRMd or MnVelRMd) have higher median large-saccade prediction errors for all three models. This is the non-tautological half of the central claim. The paper does not control for saccade amplitude. Because of the main-sequence relationship, peak and mean velocity scale with saccade size, and Fig. 2a shows prediction error increases with saccade size for all models. If subjects differ in the amplitude distribution of the saccades classified as 'large' — e.g., due to targeting bias or exclusion of blinks — then the reported rs could reflect amplitude rather than a per-subject velocity trait. The paper never reports the correlation between the velocity features and amplitude, a partial correlation, or an amplitude-matched analysis. A secondary issue is that Table 2 lacks per-subject sample counts and confidence intervals for the median errors, so the stability of the 67 subject-level points is unverified; however, the amplitude confound is the more direct threat to the stated saccade claim. The fixation-noise correlation is acknowledged by the authors as expected, so it is not the load-bearing part.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes individual differences in gaze prediction performance using three fundamentally different models (LSTM, TST, and an oculomotor plant model with Kalman filtering, OPKF) on the GazeBase random saccades dataset. It evaluates prediction error at a 40 ms prediction interval across fixation, large saccades, small saccades, and the post-saccadic critical evaluation period (CEP). The paper reports subject-to-subject variation in median prediction errors and then correlates per-subject oculomotor features with these errors. The central findings are that a fixation noise threshold is associated with poorer fixation prediction and that saccade velocity measures are associated with poorer large-saccade prediction, with the correlations consistent across all three models after Bonferroni correction.","tokens_in":11057,"tokens_out":3628,"duration_ms":32185,"significance":"If the findings hold, they provide a practical route to anticipating which subjects will be hard to predict for gaze-contingent rendering and other real-time eye-tracking applications, and they do so across fundamentally different model architectures. The study's strengths include the use of three distinct models, a relatively large test set of 67 subjects, and the application of Bonferroni correction for multiple comparisons. The cross-model consistency of the correlations is encouraging. However, the saccade-velocity result is currently threatened by a plausible amplitude confound, and the stability of the subject-level medians is not documented. These issues are addressable and should be fixed before the central claim is accepted.","major_comments":[{"comment":"The correlations between saccade-velocity features (PkVelDurRatioRMd, MnVelRMd) and median large-saccade prediction error do not control for saccade amplitude. Figure 2a shows that prediction error increases with saccade amplitude for all three models, and the main-sequence relationship implies that the velocity features also scale with amplitude. If subjects differ in the amplitude distribution of saccades classified as large, the reported rs values could reflect amplitude differences rather than a per-subject velocity trait. The authors should report the correlation between the velocity features and amplitude, a partial correlation controlling for amplitude, or an amplitude-matched analysis to support the velocity claim.","section":"4.3, Table 2"},{"comment":"The subject-level median errors used in the correlation analysis are computed without reporting the number of fixation or saccade samples per subject or confidence intervals for the medians. If some subjects contributed very few usable samples, their medians would be noisy, which could inflate or distort the Spearman correlations in Table 2. Please provide per-subject sample counts, bootstrap confidence intervals, or a minimum-sample inclusion threshold.","section":"4.2, Table 1, Figure 3"}],"minor_comments":[{"comment":"The definition of PkVelDurRatioRMd contains typographical errors ('P kV elDurRatioRM dis' and inconsistent capitalization); please correct these and clarify the units, since peak velocity divided by number of samples is not 'per millisecond' unless the sampling rate is explicitly stated.","section":"4.3.1"},{"comment":"In the Discussion, 'silency maps' should be 'saliency maps'.","section":"5"},{"comment":"The statement that code and data will be made available via a link 'to be provided in the future' should be replaced with an actual link or a clear statement of availability upon publication, as it is important for reproducibility.","section":"3.4"},{"comment":"The paper reports only the significant correlations and states that no significant correlations were found for small saccades. For transparency, please include the full set of tested correlations (or a supplement) so readers can assess the selective-reporting risk, and specify the exact number of tests used in the Bonferroni correction.","section":"4.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of CS.HC and addresses a relevant problem. The main technical concern is the amplitude confound in the saccade-velocity finding; if the authors provide adequate controls (e.g., partial correlations or amplitude matching), the paper could become acceptable. I also note that the promised data/code link should be provided before acceptance. The citation practice appears appropriate, and the appendix is helpful for assessing the robustness of the results across prediction intervals."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is honest, small-scope empirical work, and the claim that individual oculomotor features predict gaze prediction error across three very different models is worth taking seriously. But the saccade-velocity finding has a confound that needs addressing before it becomes load-bearing.\n\nWhat's new: most gaze prediction work reports aggregate error. This paper drills into per-subject medians and shows the spread is large, especially for fixations (ratios around 7x). It then correlates those medians with a small set of oculomotor features across LSTM, TST, and OPKF. The consistency of the correlations across three architectures is a genuinely useful observation, and the authors' suggestion that future papers report IQR or similar dispersion stats is sensible. Statistics are simple and appropriate (Spearman + Bonferroni), and the authors are transparent about the fixation noise result being expected.\n\nThe soft spot is in Table 2. The saccade velocity features (PkVelDurRatioRMd, MnVelRMd) are correlated with median large-saccade error. But peak velocity scales with saccade amplitude, and Fig. 2a shows error grows with amplitude. The paper doesn't control for amplitude or report whether the velocity features remain predictive after partialling out saccade size. Without that, the claim 'higher velocities → poorer prediction' may just be the known amplitude effect restated at the subject level. This is fixable with a partial correlation or amplitude-matched analysis.\n\nOther issues are secondary. Code/data are promised with a placeholder link, so nothing is reproducible as submitted. The analysis reports only significant correlations (small saccades omitted), which risks selection. No confidence intervals or per-subject sample counts for the median errors, so the stability of the 67 points is unverified. And the fixation-noise correlation is near-tautological, though the authors acknowledge it.\n\nWho it's for: people building gaze-contingent rendering or evaluating eye-tracking models; also anyone who wants a template for reporting subject-level variation. It's not a breakthrough, but it's a solid, readable empirical note.\n\nRecommendation: yes, send to review. The confound is addressable and the core observation — large subject differences and a simple per-subject predictor — likely survives. The revision should be asked to add amplitude controls, full correlation tables, and make code/data available.","headline":"Honest, small-scope study of per-subject gaze prediction; the saccade-velocity claim needs an amplitude control before it fully lands.","tokens_in":11676,"tokens_out":2797,"would_cite":true,"duration_ms":26361,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Per-subject oculomotor traits—fixation noise and saccade velocity—predict how hard a subject's gaze is to predict across three very different models.","keywords":["eye movement prediction","individual differences","fixation noise","saccade velocity","LSTM","transformer","oculomotor plant mathematical model","foveated rendering"],"falsifier":"Recompute Table 2 with per-subject medians replaced by bootstrapped medians, or with a minimum sample count per subject; if the Spearman correlations between fixation noise or saccade velocity and median error fall to near zero, the reported association was driven by noisy median estimates.","tokens_in":10644,"feed_emoji":"👁","tokens_out":5042,"duration_ms":45288,"temperature":0.7,"pith_summary":"This paper asks why the same gaze-prediction model works well for some people and poorly for others. Analysing a large eye-tracking dataset with three structurally different predictors—an LSTM, a transformer, and a Kalman-filtered oculomotor plant model—the paper finds that per-subject prediction errors vary substantially for every model. The central result is that two measurable oculomotor traits track this variation: subjects with noisier fixations have worse fixation prediction, and subjects with faster saccades have worse saccade prediction. If correct, this means a subject's eye-movement style, not just the model's architecture, sets the practical accuracy ceiling for gaze prediction, which matters for latency-sensitive foveated rendering in virtual reality.","feed_headline":"Noisy fixations and fast saccades reveal the hardest gaze predictions","feed_subtitle":"Per-subject eye-movement traits forecast prediction errors for LSTM, transformer, and Kalman-filter gaze models.","key_machinery":"The analysis is carried by per-subject profiles of median prediction error, computed separately for fixations, large saccades, and small saccades, and correlated against a small curated set of 35 radial oculomotor features plus a fixation velocity-noise threshold. The load-bearing features are Fixation Noise Threshold—the 90th percentile of radial velocity during fixations, from the MNH event-classification work—and the saccade velocity measures PkVelDurRatioRMd and MnVelRMd, which are near-duplicates. Spearman correlations with Bonferroni correction link these features to median errors for all three models. A Kendall Coefficient of Concordance appendix quantifies how similarly the models rank subjects, showing high agreement for fixations and small saccades.","core_discovery":"On the paper's own terms, the discovery is that individual oculomotor characteristics are associated with gaze prediction performance in a consistent, model-independent way. Using 322 subjects from the GazeBase random-saccade task and a 40 ms prediction interval, the authors show that for fixations the median prediction error per subject correlates strongly with a Fixation Noise Threshold (Spearman $r_s$ between 0.79 and 0.93 across the three models); the noisier a subject's fixations, the poorer the prediction. For large saccades, two nearly interchangeable velocity measures—peak velocity per sample and mean velocity per saccade—correlate with median error ($r_s$ up to 0.75), so faster saccades predict worse performance. The same oculomotor measures succeed despite the models being fundamentally different, and the subject profiles for fixations and small saccades agree strongly across models, while large-saccade profiles depend more on the model.","pith_inferences":["If the correlations generalize, a quick calibration could estimate a user's fixation noise and saccade velocity and set rendering budgets or prediction expectations per subject.","The same features might predict error on lower-quality VR eye-trackers, but noisier signals could weaken or strengthen the correlations; testing on such data would be a direct extension.","Because all three models show similar associations, the effect may be inherent to the oculomotor signal itself, suggesting that no architecture change alone will fix prediction for high-velocity or high-noise subjects."],"forward_implications":["Gaze prediction studies should report inter-subject variation, not just aggregate error, because average performance can hide subjects for whom foveated rendering fails.","Oculomotor features such as fixation noise and saccade velocity can flag difficult-to-predict subjects before or during use.","Future models could take fixation-noise and saccade-velocity measures as inputs, or be tuned to reduce subject-to-subject variation.","Prediction error ordering—fixations best, then CEP intervals, small saccades, and large saccades worst—should hold across prediction intervals, with longer intervals increasing error."],"supporting_citations":[{"why":"Supplies the GazeBase dataset, participants, and the random saccade task used for all predictions.","marker":"[33]"},{"why":"Provides the extensive set of oculomotor features from which the 35 radial fixation and saccade measures were selected.","marker":"[44]"},{"why":"Defines the MNH event classification algorithm and the Fixation Noise Threshold used as the key fixation predictor.","marker":"[42]"},{"why":"Provides the per-subject oculomotor plant model optimization procedure that the OPKF predictor relies on.","marker":"[38]"},{"why":"Supplies the transformer-based TST architecture used as one of the three prediction models.","marker":"[34]"},{"why":"Provides the lightweight LSTM architecture and the CEP evaluation approach for post-saccade prediction.","marker":"[31]"},{"why":"Earlier finding that high-velocity events are harder for gaze prediction, the prior result the saccade-velocity correlation extends.","marker":"[10]"},{"why":"Introduces eye movement prediction by Kalman filter with an integrated oculomotor plant model, the basis of OPKF.","marker":"[16]"}],"fun_headline_variants":["Eye-movement traits predict gaze prediction accuracy across models","Fixation noise and saccade speed forecast gaze prediction errors","Individual oculomotor traits determine gaze prediction performance","Noisier fixations and faster saccades worsen gaze predictions","Subject-specific eye behavior predicts gaze tracking errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that a single session's median prediction error for each subject is a stable estimate of that subject's true performance, but it does not report sample counts per subject or confidence intervals for those medians.","fun_headline_variants_meta":{"raw":{"variants":["Eye-movement traits predict gaze prediction accuracy across models","Fixation noise and saccade speed forecast gaze prediction errors","Individual oculomotor traits determine gaze prediction performance","Noisier fixations and faster saccades worsen gaze predictions","Subject-specific eye behavior predicts gaze tracking errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1282,"prompt_tokens":908,"completion_tokens":374,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":296}},"tokens_in":524,"tokens_out":374,"duration_ms":4160,"temperature":1.0,"reasoning_tokens":296,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:46:51.002668+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute Table 2 with per-subject medians replaced by bootstrapped medians, or with a minimum sample count per subject; if the Spearman correlations between fixation noise or saccade velocity and median error fall to near zero, the reported association was driven by noisy median estimates.","supporting_citations":[{"cited_title":"Gazebase, a large-scale, multi-stimulus, longitudinal eye movement dataset","cited_arxiv_id":null,"evidence_quote":"Supplies the GazeBase dataset, participants, and the random saccade task used for all predictions."},{"cited_title":"Study of an extensive set of eye movement features: Extraction methods and statistical analysis","cited_arxiv_id":null,"evidence_quote":"Provides the extensive set of oculomotor features from which the 35 radial fixation and saccade measures were selected."},{"cited_title":"A novel evaluation of two related and two independent algorithms for eye movement classification during reading","cited_arxiv_id":null,"evidence_quote":"Defines the MNH event classification algorithm and the Fixation Noise Threshold used as the key fixation predictor."},{"cited_title":"Komogortsev","cited_arxiv_id":null,"evidence_quote":"Provides the per-subject oculomotor plant model optimization procedure that the OPKF predictor relies on."},{"cited_title":"A transformer- based framework for multivariate time series representation learning","cited_arxiv_id":null,"evidence_quote":"Supplies the transformer-based TST architecture used as one of the three prediction models."},{"cited_title":"Practical perception-based evaluation of gaze prediction for gaze contingent rendering","cited_arxiv_id":null,"evidence_quote":"Provides the lightweight LSTM architecture and the CEP evaluation approach for post-saccade prediction."},{"cited_title":"Real-time gaze prediction in virtual reality","cited_arxiv_id":null,"evidence_quote":"Earlier finding that high-velocity events are harder for gaze prediction, the prior result the saccade-velocity correlation extends."},{"cited_title":"Eye movement prediction by kalman filter with integrated linear horizontal oculomotor plant mechanical model","cited_arxiv_id":null,"evidence_quote":"Introduces eye movement prediction by Kalman filter with an integrated oculomotor plant model, the basis of OPKF."}],"review_version":1}