{"id":"2f815e44-511d-46bf-aba4-5892fcccd3c8","arxiv_id":"2506.00545","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"A self-attention imputation model combined with a convolutional autoencoder refinement fills missing segments in smooth pursuit eye movements more accurately than PCHIP, SSA, and KNN, especially for long gaps.","lead":"Researchers tested a machine learning pipeline that fills in gaps in eye-tracking recordings from Parkinson's patients and healthy controls, and it reconstructed the missing data more accurately than three older methods. Because blinks and tracking losses frequently corrupt eye-movement signals, a more reliable imputation method could make smooth pursuit analysis more useful for screening and monitoring neurodegenerative disease.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Artificial missingness generation leaks held-out blink statistics (Sec. 3.5.1), making test gaps in-distribution and potentially inflating SAITS-RAE's advantage over baselines; the 'significant' improvement also lacks error bars.","rationale":"The reader's weakest_assumption correctly identifies the leakage of held-out blink statistics into the artificial missingness generation. This is the most load-bearing concern because it affects not just absolute accuracy but the relative comparison to non-learning baselines: SAITS can exploit the leaked mask distribution, whereas PCHIP, SSA, and KNN cannot. The reader's verdict of CONDITIONAL is appropriate: the paper's internal comparison is plausible but not yet convincing for real-world transfer. I add the lack of error bars as a reinforcing factor, since the margins in Table 2 are very small and the word 'significant' is used without statistical tests. However, the large-interval experiment (Table 3) shows large, likely real advantages for SAITS-RAE, so the paper should not be rejected outright. The concrete test — re-estimating blink statistics only from training participants and adding significance tests — would settle whether the central claim survives a fairer evaluation. I agree with the reader's weakest_assumption and recommend keeping the verdict CONDITIONAL until such a check is performed.","tokens_in":14411,"tokens_out":6979,"duration_ms":65520,"concrete_test":"Re-run the evaluation with a strictly held-out missingness protocol: estimate the empirical distributions of blink duration, position, and count using only the 137 training participants (excluding the 35 test participants), generate artificial gaps for the test sequences from these training-only distributions, and recompute Tables 1–3 for SAITS-RAE and all baselines. If the MAE/RMSE margins over KNN-RAE shrink to zero or change sign, the in-distribution mask generation is a confound. Additionally, report per-sequence error distributions and a paired Wilcoxon signed-rank test comparing SAITS-RAE with KNN-RAE on the current protocol; if p ≥ 0.05, the 'significant improvement' claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that SAITS-RAE significantly outperforms PCHIP, SSA, and KNN rests entirely on an evaluation with artificially inserted gaps. Section 3.5.1 states that the empirical distributions of blink duration, position, and count were 'extracted from the dataset' — explicitly 'an analysis of all sequences in the dataset', which includes the 35 held-out test participants. The artificial gaps for the test sequences are therefore drawn from the same distribution used to train SAITS. Because SAITS conditions on the missingness mask, it can learn typical gap lengths and locations (e.g., when the stimulus changes direction). The classical baselines (PCHIP, SSA, KNN) do not exploit these statistics. Thus the reported margins, especially the small ones in Table 2 (MAE 0.10 vs 0.11; RMSE 0.13 vs 0.14 against KNN-RAE), may partly reflect in-distribution masks rather than a genuinely transferable imputation advantage. Additionally, no confidence intervals, per-sequence variances, or significance tests are provided, so 'significant improvement' is not statistically supported. This is load-bearing because if the mask-distribution leak is removed or error bars are computed, the claimed superiority over the strongest classical baseline (KNN) may disappear in the standard missingness scenario, even though the large-interval results (Table 3) are more robust.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a missing-data imputation pipeline for smooth pursuit eye movement (SPEM) recordings, combining SAITS self-attention imputation on downsampled signals with cubic upsampling and a convolutional refinement autoencoder (RAE). It evaluates the pipeline on 5,504 sequences from 172 participants, comparing against PCHIP, SSA, and KNN under artificially inserted blink-like gaps and under a single 4 s large-gap condition. Time-domain and frequency-domain metrics are reported. The central claim is that SAITS-RAE significantly improves reconstruction accuracy over the compared methods while preserving spectral content.","tokens_in":14725,"tokens_out":6616,"duration_ms":61853,"significance":"If the central claim is validated, the proposed pipeline could be useful for preparing SPEM recordings for downstream biomarker extraction in Parkinson's disease studies. The paper's strengths are the use of a real clinical dataset, a transparent description of the blink-detection and artificial-missingness procedure, and the inclusion of both time- and frequency-domain metrics. However, the evidence as presented does not support the 'state of the art' language in the abstract: only three classical baselines are compared, no error bars or significance tests are reported, and the artificial test masks are sampled from blink statistics computed on the full dataset including test participants. These issues are fixable with additional experiments, and the large-interval results suggest the method has genuine potential.","major_comments":[{"comment":"The empirical distributions used to generate artificial missing values are extracted from 'all sequences in the dataset,' which includes the 35 held-out test participants. Because SAITS receives the missingness mask as an input, the test masks are drawn from the same distribution that the model saw during training, and the baseline methods cannot exploit these global blink statistics. This makes the test scenario in-distribution by construction and can inflate the SAITS-RAE advantage, particularly for the small margins in Table 2 (MAE 0.10 vs 0.11 and RMSE 0.13 vs 0.14 against KNN-RAE). Please recompute the empirical blink statistics using only training-participant data and re-report Tables 1-3, or otherwise demonstrate that the reported margins are insensitive to this choice.","section":"Section 3.5.1"},{"comment":"Tables 1-3 report single aggregate metric values with no per-sequence variability, confidence intervals, or paired significance tests. The abstract's word 'significant' is not supported by any statistical procedure. Given that several Table 2 comparisons are extremely close (e.g., MAE 0.10 vs 0.11 for SAITS-RAE vs KNN-RAE), the authors should report the distribution of per-sequence errors across the test sequences and run paired tests (e.g., Wilcoxon signed-rank) for each metric; otherwise the claimed improvements over KNN cannot be distinguished from noise.","section":"Tables 1-3"},{"comment":"The paper compares SAITS-RAE only against PCHIP, SSA, and KNN. The introduction discusses GAIN, diffusion models, and CycleGAN as related work, and the abstract claims superiority over 'other state of the art techniques.' Without at least one recent deep-learning imputation baseline (e.g., BRITS, GAIN, CSDI, or a Transformer imputer) evaluated on the same data, the state-of-the-art claim is unsupported. Please add such baselines or soften the claim accordingly.","section":"Section 4 and Abstract"}],"minor_comments":[{"comment":"The sentence 'SAITS has demonstrated state-of-the-art performance ... [11]' cites the SSSD paper [11]; the correct reference for SAITS benchmarks is the SAITS paper [26] or an appropriate benchmark study.","section":"Section 3.2"},{"comment":"The text states that 'SAITS-D achieves the lowest point-wise errors (MAE = 0.10, RMSE = 0.14)', but KNN-D also has RMSE = 0.14; please clarify the tie.","section":"Section 4.1, Table 1"},{"comment":"The histogram of blink durations lacks clear axis labels and units; please add them.","section":"Figure 6"},{"comment":"The MRE formula divides by x_i; for signals with small absolute values, MRE can be numerically unstable. Please state the threshold or method used to handle near-zero values beyond excluding exact zeros.","section":"Section 3.6.1, Eq. (2)"},{"comment":"The RAE MSE values (1.63e-3 vs 3.30e-3) are for reconstructing clean original signals, not for imputation errors. Please clarify this in the text so readers do not interpret them as imputation improvements.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"To the editor: This is an empirical benchmark paper with a fixable evaluation gap. The key concerns are the use of full-dataset blink statistics to generate test missingness and the absence of modern deep-learning baselines and statistical testing. I believe the authors can address these with additional experiments, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things before you read it. The useful result is that the long-gap robustness is real: on 4-second artificial track-loss gaps, SAITS-RAE gets MAE 0.36 against 0.88 for SSA and 1.34 for KNN. The small-gap advantage is marginal (MAE 0.10 vs 0.11 for KNN) and the 'state of the art' claim is unsupported because the comparison is only PCHIP, SSA, and KNN.\n\nWhat is genuinely new here is the combination of SAITS with a convolutional refinement autoencoder applied to smooth pursuit eye movements. That specific pipeline has not appeared in the prior literature, and the evaluation on a real Parkinson's/control dataset (5,504 sequences, 172 participants) is a plus. The long-gap experiment is the most convincing part because the gap is a fixed 4-second block, not sampled from the empirical blink distribution, so the leakage concern I will mention below does not affect that result. The RAE gives small but consistent improvements across methods, which the paper documents honestly.\n\nSoft spots. First, the abstract says 'other state of the art techniques' but no modern deep imputation baselines (GAIN, BRITS, etc.) are compared. That is an overclaim. Second, there are no error bars or significance tests. The standard-missingness margins over KNN are tiny, so 'significant improvement' is not statistically supported. Third, Section 3.5.1 says the empirical blink distributions were extracted from all sequences in the dataset, including the 35 held-out test participants. This leaks test missingness statistics into training. SAITS conditions on the missingness mask, so it can learn typical gap lengths and positions, while the classical baselines do not. This may inflate the reported advantage in the standard scenario. It is fixable by estimating distributions from the training participants only and re-running. Finally, no code or data is provided, which limits reproducibility.\n\nWho is this for? Researchers working on eye-movement analysis or physiological time-series imputation, especially for Parkinson's screening. They will find the long-gap result useful. The methodological novelty is thin for a general ML audience.\n\nMy recommendation: send it to peer review, but with the expectation of moderate revision. A referee should ask for error bars, a modern baseline or two, and a cleaner missingness protocol. If those are addressed, the long-gap result alone makes it a solid applied contribution.","headline":"A useful, incremental application paper on SPEM imputation that is worth reviewing but currently overclaims and needs error bars and a fix to its missingness generation.","tokens_in":15275,"tokens_out":3594,"would_cite":false,"duration_ms":36387,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By combining a self-attention imputer with a refinement autoencoder, this paper reconstructs missing segments of smooth pursuit eye movement recordings with lower time-domain error than PCHIP, SSA, and KNN, while keeping spectral content…","keywords":["missing data imputation","smooth pursuit eye movements","self-attention","SAITS","refinement autoencoder","blink artifacts","Parkinson's disease","biomedical time series"],"falsifier":"Use a second recording modality that does not drop out during blinks, such as a high-speed camera or a second tracker viewing the same eye, to obtain true eye position during real blinks and track losses; impute those real gaps with the SAITS-RAE pipeline and compare point-by-point against the simultaneously recorded truth. If the error on real gaps is not clearly below PCHIP, SSA, and KNN, the central claim fails.","tokens_in":14224,"feed_emoji":"👁️","tokens_out":6143,"duration_ms":52832,"temperature":0.7,"pith_summary":"The paper tries to establish that a deep-learning imputation pipeline can reconstruct missing segments in smooth pursuit eye movement recordings accurately enough to make the repaired signals usable for clinical analysis. The pipeline first imputes downsampled sequences with SAITS, a self-attention-based imputation network, then restores full resolution by cubic interpolation, and finally passes the signal through a custom convolutional autoencoder trained only on complete sequences. On 5,504 recordings from Parkinsonian patients and healthy controls, the authors report that this SAITS-RAE pipeline lowers mean absolute error, mean relative error, and root mean square error relative to PCHIP, SSA, and KNN while keeping frequency-domain error low. The advantage grows when entire four-second intervals are missing, which is the case where classical interpolation degrades most. If correct, this makes it practical to retain more eye-tracking data in studies of neurodegenerative disease instead of discarding incomplete trials.","feed_headline":"Self-attention imputation beats classical fills for eye-tracking gaps","feed_subtitle":"On smooth pursuit recordings, the pipeline preserves waveform and spectral content, even for 4-second gaps.","key_machinery":"The load-bearing object is the SAITS-RAE pipeline. SAITS (Self-Attention-based Imputation for Time Series) is a transformer-style imputer with two diagonally masked self-attention blocks, so each time step must infer its value from other time steps rather than from itself, followed by a weighted combination block that fuses the two imputation hypotheses. The RAE is a one-dimensional convolutional autoencoder with skip connections trained on complete SPEM sequences, and it refines the upsampled signal to recover fine temporal detail lost in downsampling. The evaluation machinery is equally important: artificial blinks are inserted using empirical distributions of blink duration, position, and count estimated from real recordings, and metrics are computed only at the artificially missing positions, with separate frequency-domain metrics over the whole signal.","core_discovery":"The central claim is that combining the SAITS transformer-style imputer with a refinement autoencoder yields the most accurate reconstruction of blink- and track-loss gaps in smooth pursuit eye movement sequences among the methods compared. The authors report SAITS-RAE as the global best in the time domain (MAE 0.10, RMSE 0.13, similarity 0.84 in the standard scenario) and as the strongest method when a continuous 4-second block is missing (MAE 0.36 vs 0.88 for SSA, 1.34 for KNN, 1.64 for PCHIP), while also producing the lowest errors in low-frequency spectral content. The claim is not that deep learning is always better; the paper explicitly notes that KNN often preserves the overall spectral envelope slightly better, so the argued advantage is a balanced combination of temporal fidelity, spectral preservation, and robustness to long gaps.","pith_inferences":["Because the empirical blink statistics used to generate test gaps were estimated from the entire dataset, including the held-out test participants, the reported test accuracy is likely optimistic for truly novel recording conditions; re-estimating the distributions on training participants only would be a sharper test.","The RAE is trained on complete sequences and applied uniformly, so it may pull imputed regions toward the manifold of typical SPEM shapes; whether this introduces bias for atypical or pathological signals is not addressed in the paper.","A natural next experiment the paper does not run is to check whether imputation quality changes downstream diagnostic accuracy, e.g., Parkinson's vs control classification, rather than only point-wise reconstruction error.","The deterministic nature of the pipeline means no confidence intervals are produced for imputed values; an attention-based approach could in principle be extended to output uncertainty, which would be valuable for long gaps."],"forward_implications":["Incomplete SPEM recordings no longer have to be discarded: imputed sequences can feed downstream biomarker extraction, increasing the usable sample size in studies of Parkinson's disease and other movement disorders.","The advantage over classical methods grows with gap length, so the method is aimed precisely at track-loss scenarios that defeat local interpolation.","Because frequency-domain error stays low, spectral analyses of imputed sequences, including low-frequency components below 1 Hz, remain informative.","A single model trained across all smooth pursuit tasks transfers across stimulus types without task-specific retraining, simplifying clinical deployment.","The same two-stage impute-then-refine design can be carried over to other long biomedical time series such as EEG and ECG, as the paper itself suggests."],"supporting_citations":[{"why":"Supplies the SAITS architecture, two diagonally masked self-attention blocks and a weighted combination block, which is the core imputation engine of the pipeline.","marker":"[26]"},{"why":"Source of the SPEM dataset, the twelve smooth pursuit tasks, and the blink-detection procedure used to mark and characterize missing segments.","marker":"[20]"},{"why":"Defines the PCHIP interpolation method used both as a baseline and as the upsampling step after SAITS imputation.","marker":"[19]"},{"why":"Provides the Singular Spectrum Analysis technique used as a comparison baseline for imputation.","marker":"[21]"},{"why":"Earlier application of PCHIP to smooth pursuit eye movements, establishing the prior state of the art the paper aims to surpass.","marker":"[18]"}],"fun_headline_variants":["Self-attention imputation beats conventional fills for eye-tracking gaps","Deep imputer restores smooth pursuit signals with lower errors","Attention model repairs eye-tracking gaps while preserving spectra","AI imputation of eye tracking gaps outperforms classic methods","Transformer imputer excels at missing eye data reconstruction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that gaps manufactured from the average blink statistics of this dataset behave like real blink and track-loss gaps in new recordings, so accuracy measured on these artificial gaps transfers to genuinely missing clinical data.","fun_headline_variants_meta":{"raw":{"variants":["Self-attention imputation beats conventional fills for eye-tracking gaps","Deep imputer restores smooth pursuit signals with lower errors","Attention model repairs eye-tracking gaps while preserving spectra","AI imputation of eye tracking gaps outperforms classic methods","Transformer imputer excels at missing eye data reconstruction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000489,"raw_usage":{"total_tokens":2408,"prompt_tokens":945,"completion_tokens":1463,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":1383}},"tokens_in":561,"tokens_out":1463,"duration_ms":9575,"temperature":1.0,"reasoning_tokens":1383,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:02:16.571492+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use a second recording modality that does not drop out during blinks, such as a high-speed camera or a second tracker viewing the same eye, to obtain true eye position during real blinks and track losses; impute those real gaps with the SAITS-RAE pipeline and compare point-by-point against the simultaneously recorded truth. If the error on real gaps is not clearly below PCHIP, SSA, and KNN, the central claim fails.","supporting_citations":[{"cited_title":"SAITS: Self-attention-based imputation for time series","cited_arxiv_id":null,"evidence_quote":"Supplies the SAITS architecture, two diagonally masked self-attention blocks and a weighted combination block, which is the core imputation engine of the pipeline."},{"cited_title":"Estimation of the cyclopean eye from binocular smooth pursuit tests","cited_arxiv_id":null,"evidence_quote":"Source of the SPEM dataset, the twelve smooth pursuit tasks, and the blink-detection procedure used to mark and characterize missing segments."},{"cited_title":"Monotone piecewise cubic interpolation","cited_arxiv_id":null,"evidence_quote":"Defines the PCHIP interpolation method used both as a baseline and as the upsampling step after SAITS imputation."},{"cited_title":"Analysis of time series structure: SSA and related techniques","cited_arxiv_id":null,"evidence_quote":"Provides the Singular Spectrum Analysis technique used as a comparison baseline for imputation."},{"cited_title":"Baseline wander removal applied to smooth pursuit eye movements from parkinsonian patients","cited_arxiv_id":null,"evidence_quote":"Earlier application of PCHIP to smooth pursuit eye movements, establishing the prior state of the art the paper aims to surpass."}],"review_version":1}