{"id":"39289724-d7f6-4eb4-9337-ffae2cc18778","arxiv_id":"2506.07488","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new open dataset combines EEG, eye-tracking, and high-speed video across motor imagery, motor execution, SSVEP, and P300 BCI paradigms from 31 subjects.","lead":"This paper presents a large multimodal dataset of EEG, eye-tracking, and high-speed video recordings from 31 subjects across four BCI paradigms, totaling over 46 hours. It provides a synchronized resource for studying blinks and eye movements as both artifacts and potential control or feature signals in brain-computer interfaces.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The video-derived blink ground truth assumes symmetric blinking between the eyes, and the cross-modal validation in Fig. 7 depends on this unvalidated assumption; the Blinks column should not be treated as ground truth until the assumption is checked.","rationale":"The reader's weakest assumption is the same one I would stress: it is the only stated assumption that directly underwrites a derived, merged product (the Blinks column) that the paper then uses to validate the dataset's cross-modal consistency. The central claim that raw simultaneous EEG, eye-tracking, and video exist is supported by the device descriptions, synchronization flow, and missing-data accounts, all of which are internally coherent. Other issues—such as the unreadable SNR formula in Eq. (1) and the impossible confidence interval for S02 in Table 4 (89% outside [69%; 79%])—are presentation-level errors that are correctable and do not threaten the existence or availability of the raw data. The symmetric-blinking concern, in contrast, affects the validity of a key derived label and the Fig. 7 cross-modal validation, so it is the most load-bearing technical risk. I did not independently access or execute the released code and data, so the appropriate resolution is a concrete empirical check rather than a change of the conditional verdict.","tokens_in":17479,"tokens_out":9337,"duration_ms":108435,"concrete_test":"From the released Synapse repository, download a stratified subset (e.g., 5 subjects × 1 session each, spanning paradigms). Run the provided video-processing code to obtain left-eye eyelid-position/blink events, and independently detect blink events from the right-eye EMG channel (and vertical EOG) using an amplitude threshold with a 20 ms refractory window. Compute onset differences for matched events: if the median absolute onset difference exceeds one video frame (6.7 ms), or if more than 5% of video blinks have no EMG match within ±30 ms, the symmetric-blinking assumption fails for those sessions. Report the same confusion metric separately for partial versus full blinks, and if the assumption fails, relabel the Blinks column as left-eye-only rather than a general blink reference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Methods section (Multimodal acquisition, p.7) states that the high-speed camera records only the left eye because the EOG electrodes on that side are positioned farther from the eyelids, and that this setup 'allows for the extraction of eyelid position from the video, operating under the assumption of symmetric blinking between both eyes.' The Blinks column merged into the EEG files (Data records, Table 2) is evidently derived from this single-eye video, and the technical validation (Fig. 7) uses it to claim that right-eye EMG and frontopolar EEG 'record similar patterns, which are in alignment with the eyelid movement data captured in the video recordings.' If blinking is not consistently symmetric—for example, in partial blinks or lid-lag producing onset/offset differences between the eyes—then the video-derived blink labels will misalign with the right-eye EMG channel, and downstream analyses that use the Blinks column as reference will inherit that bias. The paper states the assumption but offers no quantitative check, no exclusion of asymmetric events, and no sensitivity analysis. This is load-bearing for the dataset's core purpose of multimodal ocular-activity analysis, although it does not affect the raw availability of the three independent streams.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a multimodal dataset combining 65-channel EEG, Tobii TX300 eye-tracking, and high-speed video (Phantom M310) recorded from 31 participants across 63 sessions, with four BCI paradigms (motor imagery, motor execution, SSVEP, and P300 speller). The manuscript documents the acquisition setup, device synchronization via E-Prime, Arduino, and a light sensor, preprocessing steps, missing-data statistics, and technical validation including blink metrics, SNR/ERD/ERS, ERP, SSVEP, pupil analyses, and MI classification with the authors' ABCD algorithm compared with ICA and ASR. The data and code are made openly available online.","tokens_in":17711,"tokens_out":5498,"duration_ms":67384,"significance":"The dataset addresses a real gap in BCI research by providing synchronized EEG, eye-tracking, and high-speed video across multiple paradigms, which is valuable for studying ocular artifacts and developing robust artifact-correction methods. The paper is strong in its explicit synchronization protocol, detailed missing-data accounting, and inclusion of validation analyses with confidence intervals. The central claim—that a large multimodal dataset is available—is largely supported by the raw data availability. However, the reliability of the derived Blinks column, which underpins the cross-modal validation, rests on an unvalidated symmetry assumption, and several documentation errors weaken the paper's usability claims.","major_comments":[{"comment":"The high-speed camera records only the left eye, and the Blinks column (Data records, Table 2) is used in Fig. 7 to argue that right-eye EMG/EOG and frontopolar EEG 'record similar patterns... in alignment with the eyelid movement data captured in the video recordings.' The symmetric-blinking assumption is stated but never quantitatively checked; partial blinks and lid lag can produce asymmetric onset/offset between the eyes. Because the Blinks column is a derived feature, this issue is load-bearing for the dataset's multimodal ocular-analysis purpose. Please provide a symmetry validation (e.g., a subset of binocular video recordings or a comparison of left-eye video with right-eye EOG/EMG) or explicitly scope the Blinks column as left-eye-only and soften the cross-modal validation claims.","section":"Multimodal acquisition (p.7) and Fig. 7"},{"comment":"The sentence listing eye-tracking trials that lack a common time reference for subject S06's first session includes 'MI241 and P3005L261,' but these identifiers correspond to S24 and S26, respectively, not to S06. This inconsistency makes the missing-data documentation unreliable for those files. Please correct the list and verify it against the public repository.","section":"Missing data section"},{"comment":"The manuscript does not specify whether the Blinks column was generated from EEG signals (via the ABCD algorithm) or from high-speed video eyelid tracking. Table 2 shows Blink=1 at a time when FP1 amplitude is low (17.19 μV), suggesting video-derived timing, but the text in the Preprocessing section says 'Blinks are identified through the methodology outlined in [21]' (an EEG-based method). Downstream users will treat this column as ground truth, so the source and exact parameters must be stated explicitly.","section":"Data records, Table 2, and Preprocessing"},{"comment":"Several confidence intervals in Table 4 are impossible because they do not contain the point estimate; for example, S02 is listed as 89% [69%; 79%]. This indicates transcription errors in a key validation table. All entries should be rechecked and corrected.","section":"Table 4"}],"minor_comments":[{"comment":"There are numerous typos and formatting errors, including 'opEN' in the running footer, 'Ver y' in Table 1 and the text, 'Edimburg' for 'Edinburgh', and a duplicated 'Multimodal acquisition' heading.","section":"Throughout"},{"comment":"The section heading 'SNr plots and data quality validation' uses an inconsistent abbreviation; it should be 'SNR' for clarity.","section":"SNr plots section heading"},{"comment":"The claim that the dataset 'uniquely provides simultaneous electrophysiological recordings, video capture, and synchronized eye-tracking' would be better supported by citing and comparing with existing EEG+eye-tracking datasets and specifying exactly which combination is new.","section":"Background & Summary"},{"comment":"The units for eye correction ('K dioptre') and the meaning of the 'Decile' column are not defined; please add a brief explanation in the table caption or methods.","section":"Table 1"},{"comment":"This section is lengthy and relies on the authors' prior methods; a shorter summary focused on the chosen 63-session target would be more appropriate for a data descriptor.","section":"a priori sample size estimation"}],"recommendation":"major_revision","confidential_remarks":"The dataset appears valuable and the raw streams are likely reusable, but the unvalidated symmetry assumption for the Blinks column and the missing-data/CI documentation errors are substantive. The paper could become acceptable for Scientific Data after the authors add a symmetry check or scope the Blinks column, correct the missing-data list and Table 4, and clarify the Blinks column's provenance. The overbroad 'unique' claim should also be tempered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a data descriptor for a genuinely useful resource: synchronized EEG, eye-tracking, high-speed video, and questionnaire data from 31 subjects across four BCI paradigms (MI, ME, SSVEP, P300). To my knowledge, no prior public dataset combines all four with video of the eye at this scale. The data and code are online, CC0, and the paper documents acquisition, synchronization, and preprocessing carefully.\n\nWhat it does well: the missing-data reporting is honest (e.g., 6.5% of eye-tracking trials lack a common time base; specific videos missing). The validation is reasonable: blink statistics per subject, ERD/ERS, ERP, SSVEP, and classification with Clopper-Pearson CIs. They also include demographic and facial-landmark info, which supports secondary analyses. The synchronization procedure via E-Prime, Arduino, and Cedrus is described in enough detail to reimplement.\n\nThe main concern is the symmetric-blinking assumption. Video records only the left eye, and the Blinks column in the EEG files is derived from that video, assuming both eyes move together. The paper states this but offers no quantitative check (e.g., comparing left-eye video against right-eye EMG for a subset). If asymmetric blinks are frequent, the Blinks column will mislabel some right-eye events. That is a genuine limitation for anyone using the Blinks column as ground truth. However, it does not affect the raw EEG, eye-tracking, and video streams themselves, which are the core deliverable. A user can compute their own blink labels from the EOG/EMG channels or right-eye video if they care. The paper should have validated this, but it's a fixable gap, not a fatal flaw.\n\nTwo smaller items: Eq. (1) SNR formula is unreadable as typeset—should be corrected or omitted. And the validation leans on the authors' own blink-correction algorithm (ref 21), which outperforms ICA/ASR in their Table 4; that's fine as a demonstration but should not be read as independent benchmarking.\n\nBottom line: the dataset is the contribution, and it appears solid. The central claim holds up. The paper deserves a serious referee, and I'd send it out. I'd cite it if I worked on BCI artifact removal or ocular-EEG integration.","headline":"A genuinely useful open multimodal BCI dataset with a few documented limitations; the symmetric-blinking assumption deserves a check, but the raw streams are the real deliverable.","tokens_in":18202,"tokens_out":2559,"would_cite":true,"duration_ms":28180,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"New multimodal dataset combines EEG, eye-tracking, and video to study blinks across four BCI paradigms.","keywords":["EEG","eye-tracking","high-speed video","blink detection","brain-computer interface","motor imagery","SSVEP","P300"],"falsifier":"If a subject is recorded with separate bilateral video and EEG/EOG/EMG, and a substantial fraction of blinks show timing or amplitude differences between the two eyes (beyond a small tolerance), the single-eye video blink labels would misalign with right-eye signals, indicating that the dataset's cross-modal labels could be biased.","tokens_in":17296,"feed_emoji":"👁️","tokens_out":3414,"duration_ms":37384,"temperature":0.7,"pith_summary":"This paper introduces a publicly available multimodal dataset that records electroencephalogram (EEG) brain signals, eye-tracking gaze data, and high-speed video of the left eye simultaneously, from 31 healthy volunteers across 63 sessions and four brain-computer interface (BCI) tasks. The authors' central claim is that this combination, together with questionnaires on each participant's mental state and physical characteristics, is unique in letting researchers study blinks and other eye movements with independent visual, electrical, and gaze-based evidence. Because blinks can either corrupt EEG or serve as intentional control signals, the dataset is meant to support algorithms that handle eye-induced artifacts and improve task classification, and to test how well BCI methods generalize across paradigms. The paper also reports quality checks, including per-subject blink statistics, signal-to-noise plots, and motor-imagery classification accuracies.","feed_headline":"New multimodal BCI dataset tracks blinks in EEG, video, and gaze","feed_subtitle":"Over 46 hours of synchronized recordings across four paradigms let researchers study blinks as noise or as control signals.","key_machinery":"The central mechanism is the synchronized multimodal acquisition pipeline. An Arduino Nano converts E-Prime trigger codes into square waves that regulate the Phantom camera shutter, while the same triggers go to the EEG system; a light sensor detected by the Tobii eye-tracker provides the common time reference, so every EEG sample, gaze point, and video frame can be compared. Blink detection relies on computer-vision tracking of eyelid landmarks in the video, cross-validated against EEG and EOG/EMG signals.","core_discovery":"The core contribution is the dataset itself: over 46 hours of recordings from 31 subjects, with 2,520 trials each of motor imagery, motor execution, and steady-state visually evoked potentials, and 5,670 P300 trials, all aligned to a common time base. EEG was sampled at 1000 Hz from 65 channels, the eye-tracker at 300 Hz, and the high-speed camera at 150 frames per second, with triggers from E-Prime synchronized to all three devices. Blink labels are provided in the EEG files, derived from video-based eyelid tracking, and the authors demonstrate inter-modal consistency by showing that frontopolar EEG, EOG, EMG, and video-derived eyelid movements align during blinks. The dataset is released under CC0 with code for data loading and reproduction.","pith_inferences":["The single-eye video assumption could be empirically tested using the bilateral EOG/EMG and gaze data already in the dataset; if asymmetric blinks are common, future versions should record both eyes.","Because the high-speed camera stores only about seven minutes of video per run, the dataset cannot capture long-term blink changes; continuous webcam recordings could extend this coverage.","The facial-landmark and blink-width distributions might support new biometric identification studies, although the paper's anonymization claims would need scrutiny.","The acquisition design could be replicated with consumer hardware, such as a webcam and a budget EEG system, to test whether video-based blink ground truth remains reliable without a high-speed camera."],"forward_implications":["Researchers can train artifact-correction algorithms that exploit video-verified blink ground truth rather than EEG-only heuristics.","Because the same participants performed four paradigms, the dataset enables cross-paradigm generalization tests for BCI classifiers.","The per-subject blink statistics support studies of inter- and intra-subject variability in blink amplitude, width, and frequency.","The pupil-size and gaze data can be used to study cognitive load and attention differences across BCI tasks.","The dataset provides a benchmark for comparing blink correction methods such as ICA, ASR, and the authors' ABCD algorithm."],"supporting_citations":[{"why":"Provides the ABCD blink detection and correction algorithm used to identify blinks in EEG and to compute artifact-corrected classification accuracies.","marker":"[21]"},{"why":"The Synapse repository that hosts the multimodal dataset and is the primary deliverable of the paper.","marker":"[22]"},{"why":"The a priori sample size determination method used to justify the 63-session, 46-hour recording plan.","marker":"[14]"},{"why":"The iBUG 300-W face dataset used to train the facial landmark detector that extracts subject-specific face measures.","marker":"[17]"},{"why":"The original P300 speller paradigm that the P300 tasks are based on.","marker":"[18]"},{"why":"The checkerboard stimulus paradigm used to generate P300 flash sequences with reduced adjacent-letter confusions.","marker":"[19]"},{"why":"Independent Component Analysis, one of the artifact removal methods compared in the classification accuracy validation.","marker":"[32]"},{"why":"Artifact Subspace Reconstruction, another artifact removal method compared in the classification accuracy validation.","marker":"[33]"}],"fun_headline_variants":["46-hour multimodal BCI dataset captures blinks across four paradigms","EEG, eye-tracking, video combined in blink-focused BCI dataset","New dataset links EEG, gaze, video for BCI blink studies","Blink-rich dataset spans four BCI paradigms, 31 subjects"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Blinking is assumed to be symmetric between the two eyes, which justifies recording only the left eye on video and using those eyelid positions to label blinks in the right-eye EOG/EMG electrodes.","fun_headline_variants_meta":{"raw":{"variants":["46-hour multimodal BCI dataset captures blinks across four paradigms","EEG, eye-tracking, video combined in blink-focused BCI dataset","New dataset links EEG, gaze, video for BCI blink studies","Blink-rich dataset spans four BCI paradigms, 31 subjects"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000176,"raw_usage":{"total_tokens":1277,"prompt_tokens":922,"completion_tokens":355,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":279}},"tokens_in":538,"tokens_out":355,"duration_ms":4407,"temperature":1.0,"reasoning_tokens":279,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:31:27.365826+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a subject is recorded with separate bilateral video and EEG/EOG/EMG, and a substantial fraction of blinks show timing or amplitude differences between the two eyes (beyond a small tolerance), the single-eye video blink labels would misalign with right-eye signals, indicating that the dataset's cross-modal labels could be biased.","supporting_citations":[{"cited_title":"& Pantic, M","cited_arxiv_id":null,"evidence_quote":"The iBUG 300-W face dataset used to train the facial landmark detector that extracts subject-specific face measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The checkerboard stimulus paradigm used to generate P300 flash sequences with reduced adjacent-letter confusions."},{"cited_title":"& Sejnowski, T","cited_arxiv_id":null,"evidence_quote":"Independent Component Analysis, one of the artifact removal methods compared in the classification accuracy validation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Artifact Subspace Reconstruction, another artifact removal method compared in the classification accuracy validation."}],"review_version":1}