{"id":"fbc13dff-cfb3-407a-af1e-6eee2130382e","arxiv_id":"2412.04888","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new multimodal driving-simulator dataset combining facial video, steering-wheel ECG, eye tracking, and vehicle signals for fatigue and takeover studies.","lead":"The paper introduces VTD, a driving-simulator dataset with 600 minutes of fatigue driving data from 15 subjects, 102 takeover scenarios from 17 drivers, and synchronized video, ECG, eye-tracking, and vehicle signals. It is a candidate resource for training systems that monitor tired or distracted drivers in human-machine co-driving, but the dataset itself is not currently linked or released.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"VTD's central claim as a usable benchmark cannot be verified because the paper provides no dataset release, repository, DOI, or access mechanism; the resource exists only as a self-report.","rationale":"The reader's verdict is CONDITIONAL and I do not move it. The decisive issue is data availability: the paper claims a usable, standardized benchmark but supplies no access path, so the existence, synchronization, and annotations cannot be checked. This precedes the reader's stated weakest assumption about subjective fatigue labels, because even a well-designed label protocol cannot be audited without the data. That said, the reader's rationale already mentions the absence of release, so there is partial agreement even though the formal weakest_assumption field emphasizes label subjectivity. I would keep the verdict CONDITIONAL: release the data with documentation, add a label-reliability analysis (e.g., inter-rater agreement and distribution), and include baseline experiments or signal-quality validation; then the dataset claim could be verified. I do not see a need to escalate to REJECT because the protocol is described in enough detail that a conditional acceptance as a dataset proposal is appropriate.","tokens_in":11125,"tokens_out":6172,"duration_ms":61366,"concrete_test":"Search the arXiv source, the authors' institutional pages, and public research-data repositories (Zenodo, figshare, GitHub) for any VTD download link or DOI. If none exists, contact the corresponding author and request the dataset or a minimal file manifest with synchronized modality files and labels. The concern lands if no accessible release is provided within a reasonable window; it is resolved if a link to documented, synchronized data appears.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim is that VTD is a comprehensive, synchronized, large-scale benchmark (Abstract; Section IV.A). For a dataset paper, the minimal condition for this claim to hold is that the dataset can actually be obtained and used. The manuscript contains no Data Availability statement, no download link, no repository URL, no DOI, and no file manifest; Section IV only describes contents and applications. Consequently, the claimed 600 minutes of fatigue data and 102 takeover experiments cannot be inspected, downloaded, or benchmarked against. This also blocks independent checks that would be routine once the data exist: verifying synchronization across frontal video, ECG, eye tracking, and vehicle signals; auditing the Table III ANOVA features for label leakage or subject dependence; and assessing label reliability. The problem is not an internal inconsistency; it is that the central artifact is asserted rather than supplied. Without access, the 'standardized platform for benchmarking' claim is unverifiable, regardless of the statistical results reported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents VTD, a multimodal dataset for driver state and behavior perception, collected in a driving simulator. For the fatigue portion, 15 participants drove for 40 minutes each under four pre-fatigue states while frontal RGB video, ECG from steering-wheel electrodes, eye tracking, and vehicle signals were recorded; labels were derived from KSS/SSS self-reports combined with staff observation. For the distraction/takeover portion, 17 participants performed 102 takeovers under visual and auditory subtasks across three road scenarios. The manuscript describes the data collection infrastructure, feature extraction (Mediapipe landmarks, EAR/MAR, head pose, heart-rate variability), a one-way ANOVA of time-series features against fatigue levels, and a comparison table against existing datasets. The intended contribution is that VTD is a comprehensive, synchronized, large-scale benchmark for cross-modal driver perception.","tokens_in":11231,"tokens_out":6095,"duration_ms":62210,"significance":"If the data were actually released and the labeling protocol validated, VTD would be a valuable addition to driver monitoring datasets, chiefly because the contactless steering-wheel ECG channel is combined with synchronized frontal video, eye tracking, and vehicle signals. The four-level pre-fatigue protocol and the repeated five-minute self/staff reporting cadence are a reasonable experimental design, and the paper usefully situates the work against existing DMS datasets. However, the benchmark claim cannot be assessed from the manuscript as it stands: the data are not accessible, no baseline experiments are run, the statistical evidence in Table III has internal inconsistencies, and the fatigue labels carry a circularity risk. The paper deserves a major revision rather than acceptance in its current form.","major_comments":[{"comment":"The central claim that VTD is 'a comprehensive, well-structured, large-scale dataset' and a 'standardized platform for benchmarking' cannot be verified because the manuscript supplies no dataset release mechanism: there is no Data Availability statement, repository URL, DOI, download link, file manifest, sample data, or annotation example anywhere in Sections III–V. For a dataset paper, accessibility is the minimal condition for the benchmark claim; without it, the claimed 600 minutes of fatigue data and 102 takeover experiments remain a self-report. Please add a persistent public release (or a documented controlled-access procedure with an institutional contact), a file-format and folder-structure description, and at least one sample snippet of each modality, including the annotation format.","section":"IV.A and Abstract"},{"comment":"The fatigue labels are constructed from KSS/SSS self-reports and staff observation, and the same types of behavioral indicators (eye closure, head pose, facial action) are then tested in the Table III ANOVA. To the extent that the staff used those very signals to assign fatigue levels, the significant F statistics may reflect the labeling rule rather than an independent physiological correlate. The manuscript does not report inter-rater reliability, the number of label disagreements, the label distribution across the four pre-fatigue states, or any validation of the self-reports against objective performance. Please report these quantities and specify exactly how self-reports and staff observations were combined into the final labels.","section":"III.B and III.E"},{"comment":"The statistical evidence in Table III is internally inconsistent and therefore not load-bearing. The caption defines '++++' as α<0.01 and '+++' as 0.01≤α<0.05, but the Steering Angle(SD) row reports P=5.5569×10−2 ≈ 0.056, which is not significant at α=0.05 yet is marked '+++', while the Pedal(SD) row reports P=1.4349×10−2 and is marked '++++' even though it should be '+++' by the caption; the P value for Steering Angle(SD) also looks like a transcription of the Blinking Rate P value. In addition, the text alternates between '11-dimensional' (Section III.E) and '10-dimensional' (Table III, Table IV) time series. Please recompute the table, state the exact hypothesis and multiple-comparison correction, and confirm whether samples from the same driver are treated as independent in the ANOVA.","section":"Table III"},{"comment":"The benchmark claim is unsupported by any experimental demonstration. Section III.E reports a 4:1:1 train/validation/test split and Table III gives univariate ANOVAs, but no classification or regression baseline is run on VTD, no metric (accuracy, F1, AUC) is reported, and no comparison with existing datasets such as DMD or FatigueView is made under a common protocol. To substantiate the 'benchmarking platform' contribution, add at least one well-defined baseline task (e.g., fatigue-level classification and takeover reaction-time prediction) with subject-independent evaluation and standard metrics.","section":"IV.A and Section III.E"}],"minor_comments":[{"comment":"The Abstract says '600 minutes' while Section IV.A says '10-hour fatigue driving data' and Table IV lists '630 min videos' for VTD; clarify whether the 630 minutes includes the takeover videos and whether the 600 minutes refers only to the fatigue protocol.","section":"Abstract and Table IV"},{"comment":"Section III.C states that '34 groups of visual and auditory subtasks' were established but Table IV reports '6 types of scenarios, 102 takeover experiments'; specify how the 34 groups and 6 scenario types relate to the three conditions (straight path, roundabout cut-in, roundabout obstacle avoidance).","section":"III.C"},{"comment":"The sentence 'The dataset labels were modified based on KSS and SSS fatigue scales' is vague; state which scale was collected at which time point and how the final ordinal label was computed.","section":"III.B"},{"comment":"The relation between 600 minutes of fatigue recording, 60-second time slices, and 480 valid samples is unexplained; specify the windowing, overlap, and exclusion criteria used to obtain 480 samples.","section":"III.B"},{"comment":"The landmark indices P82, P87, P312, etc. are undefined; give a pointer to the canonical face model or include a landmark index diagram so that the MAR and EAR definitions are reproducible.","section":"Equations (1)–(2)"}],"recommendation":"major_revision","confidential_remarks":"The main issue for the editor is that the paper's core artifact is inaccessible: the authors describe a dataset but provide no way to obtain, inspect, or benchmark against it. If the data cannot be made available (publicly or through a clear controlled-access procedure) within the revision, the paper should not be accepted as a dataset paper. The statistical inconsistencies in Table III and the label-circularity concern should also be addressed head-on; the current text gives no indication that these limitations exist."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: VTD is a genuinely new combination—frontal video, steering-wheel contact ECG, eye tracking, and vehicle signals under defined fatigue and takeover protocols in a simulator—and the protocol write-up is detailed enough to credit. But the paper's central claim that VTD is a usable benchmark cannot be checked: there is no dataset link, no repository, no sample frames, no annotation examples, and no baseline experiments. As written, it is a dataset proposal without the dataset.\n\nWhat's good: The combination is not in any of the cited alternatives. The pre-fatigue protocols (A1-A4) and the takeover scenario definitions (straight/roundabout cut-in/obstacle avoidance, visual/auditory subtasks) are concrete and reproducible in principle. The feature table (Table III) gives a first look at which time-series features separate fatigue levels, and the MAR/EAR formulas and Mediapipe pipeline are usable. The authors also state the intended split (4:1:1) and the 60-second windowing, so evaluation would be standardized if the files existed.\n\nSoft spots, in rough order: (1) No data release. This is load-bearing. No DOI or URL; the claim 'standardized platform for benchmarking' cannot be assessed. (2) Labels come from KSS/SSS self-reports plus staff observation, with no inter-rater reliability or validation against an objective measure. Since staff observation likely uses the same behavioral cues (eye closure, head tilt, etc.) that later become features, the significant ANOVA results in Table III could partly reflect labeling circularity. (3) Inconsistency: Table IV says 630 min videos while the abstract says 600 min. (4) 'Piezoelectric tactile data' overstates the steering-wheel ECG electrodes; they are contact ECG, not piezoelectric tactile sensing. That matters because 'visual-tactile fusion' is a headline contribution. (5) Sample sizes (15/17, simulator) are modest; the authors should be less free with 'large-scale.'\n\nNone of these are fatal if the authors treat this as a proposal. But as a paper claiming an existing, comprehensive benchmark, it needs substantial revision: release the data (or a curated sample), add signal-quality metrics, baseline models, and label agreement, and reconcile the counts.\n\nMy recommendation: send it to peer review with the clear expectation of major revision—a serious referee could verify protocol details and push for data release. I would not cite it yet, but I'd want to see the next version.","headline":"A promising dataset proposal whose central artifact—the data—is not actually available; the protocol detail is real, but the benchmark claim cannot be evaluated until VTD is released.","tokens_in":11785,"tokens_out":2327,"would_cite":false,"duration_ms":23948,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"VTD is a new multimodal benchmark that synchronizes visual driver monitoring (frontal video, eye tracking) with tactile sensing (ECG from steering-wheel electrodes) and vehicle signals, built on 10 hours of induced-fatigue driving from 15…","keywords":["visual and tactile data","driver state","driver behavior","intelligent cockpit","autonomous vehicles","driver fatigue","takeover scenarios","multimodal dataset"],"falsifier":"If the released dataset's modality timestamps are not aligned within the stated frame rates (e.g., video at 30 fps versus ECG at 250 Hz with drift exceeding 100 ms), or if the KSS and SSS labels show no agreement with objective measures such as PERCLOS or heart-rate variability, the central claim of a synchronized, reliable benchmark would be undercut.","tokens_in":10896,"feed_emoji":"🚗","tokens_out":4005,"duration_ms":40083,"temperature":0.7,"pith_summary":"The paper claims to establish VTD, a multimodal dataset for studying driver fatigue and distraction in human-vehicle co-driving. It synchronizes frontal video, eye tracking, ECG acquired through steering-wheel electrodes, and vehicle signals, collected from 15 fatigued drivers (10 hours) and 17 distracted drivers (102 takeover trials). The central assertion is that this visual-tactile combination, with human-in-the-loop subjective labeling, forms a standardized benchmark for cross-modal driver state and behavior perception. A sympathetic reader would care because current public datasets are mostly single-modality and lack synchronized physiological and behavioral signals.","feed_headline":"Driver dataset fuses video, ECG, and vehicle signals","feed_subtitle":"Ten hours of fatigue driving and 102 takeover trials give AI a multimodal benchmark for driver monitoring.","key_machinery":"The central mechanism is the dataset itself and its synchronized multi-modal acquisition platform: a driving simulator (Logitech G29) with flexible ECG electrodes built into the steering wheel, an RGB camera, Tobii Glasses3 eye tracking, and vehicle signal logging. The fatigue induction protocol defines four pre-fatigue states (A1 regular sleep, A2 nap deprivation, A3 partial night sleep deprivation, A4 total night sleep deprivation) and labels fatigue via KSS and SSS self-reports combined with staff observation. The takeover protocol defines three driving conditions and visual and auditory subtasks, measuring reaction time as the period from the takeover request to both hands returning to the wheel, and execution time from steering wheel and pedal thresholds.","core_discovery":"On the paper's own terms, VTD is a new large-scale multimodal dataset that pairs visual signals (RGB facial video, eye tracking, head pose) with tactile and physiological signals (ECG via steering-wheel electrodes, vehicle control signals) under controlled fatigue and distraction protocols. It provides an 11-dimensional time series for fatigue, including EAR, PERCLOS, blinking rate, MAR, head tilt, R-R intervals, SDNN, RMSSD, steering angle, pedal, speed, and transverse angular velocity, plus takeover reaction and execution times. The paper argues that this fills the gap left by single-mode datasets and enables cross-modal fusion algorithms for driver monitoring.","pith_inferences":["The steering-wheel ECG approach, if validated in real vehicles, could enable continuous driver monitoring without wearable sensors, but the simulator environment may not capture the motion artifacts and electrical noise of real-road driving.","Because fatigue labels rely on subjective KSS/SSS ratings and staff observation, benchmark users may need robust or semi-supervised methods to handle label noise, a direction the paper does not explore.","The combination of eye tracking and ECG could be extended beyond fatigue to estimate cognitive workload or engagement in non-driving tasks, which the paper mentions only through the QN-ACTR load-rate model.","Findings from this simulator-based dataset may transfer to real driving only after domain adaptation; the paper does not provide real-world validation, so users should treat simulator-to-real generalization as an open question."],"forward_implications":["Researchers can train cross-modal driver monitoring models that fuse facial video with ECG and vehicle signals using VTD's synchronized time series.","The 102 takeover scenarios with defined reaction and execution times allow benchmarking of takeover performance under visual versus auditory subtasks across different road geometries.","The steering-wheel ECG collection method offers a less intrusive alternative to wearable ECG sensors for driver monitoring, potentially transferable to real vehicles.","The 10 hours of fatigue data with 11-dimensional time series and a fixed 4:1:1 train-validation-test split support the development and comparison of fatigue detection algorithms.","VTD's multimodal design enables investigation of how visual features (e.g., PERCLOS, head tilt) and physiological features (e.g., SDNN, RMSSD) jointly indicate fatigue levels, going beyond single-modality approaches."],"supporting_citations":[{"why":"Supplies the Mediapipe Facemesh model used to extract 478 facial keypoints for computing MAR and EAR values in the fatigue dataset.","marker":"[33]"},{"why":"3MDAD is a public multimodal distraction dataset that VTD compares against, positioning VTD as more comprehensive in scenarios and modalities.","marker":"[15]"},{"why":"DMD is a large-scale multi-modal driver monitoring dataset used as a comparison point for comprehensive fatigue and distraction data.","marker":"[16]"},{"why":"YawDD represents a single-modality yawning detection dataset that VTD aims to surpass with multi-modal synchronized data.","marker":"[25]"},{"why":"RLDD provides a realistic drowsiness detection baseline dataset that VTD contrasts with its own controlled fatigue protocol.","marker":"[28]"},{"why":"FatigueView is a multi-camera drowsiness detection dataset cited to support the need for richer contextual data beyond single-modality visual features.","marker":"[24]"}],"fun_headline_variants":["Multimodal driver data: 10h fatigue, 102 takeovers","VTD dataset blends face, ECG, and vehicle signals","Driver state benchmark: eye, heart, and steering","Fatigue and distraction: 17 drivers, cross-modal signals","New dataset fuses video, tactile, and vehicle data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The fatigue labels come from KSS and SSS self-reports and staff observation rather than any objective physiological or behavioral ground truth, so the benchmark's value depends on those subjective ratings faithfully tracking true fatigue.","fun_headline_variants_meta":{"raw":{"variants":["Multimodal driver data: 10h fatigue, 102 takeovers","VTD dataset blends face, ECG, and vehicle signals","Driver state benchmark: eye, heart, and steering","Fatigue and distraction: 17 drivers, cross-modal signals","New dataset fuses video, tactile, and vehicle data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00026,"raw_usage":{"total_tokens":1506,"prompt_tokens":777,"completion_tokens":729,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":393,"completion_tokens_details":{"reasoning_tokens":644}},"tokens_in":393,"tokens_out":729,"duration_ms":7131,"temperature":1.0,"reasoning_tokens":644,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:10:09.942166+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If the released dataset's modality timestamps are not aligned within the stated frame rates (e.g., video at 30 fps versus ECG at 250 Hz with drift exceeding 100 ms), or if the KSS and SSS labels show no agreement with objective measures such as PERCLOS or heart-rate variability, the central claim of a synchronized, reliable benchmark would be undercut.","supporting_citations":[{"cited_title":"Modeling drowsy driving behaviors,","cited_arxiv_id":null,"evidence_quote":"Supplies the Mediapipe Facemesh model used to extract 478 facial keypoints for computing MAR and EAR values in the fatigue dataset."},{"cited_title":"A novel public dataset for multimodal multiview and multi- spectral driver distraction analysis: 3mdad,","cited_arxiv_id":null,"evidence_quote":"3MDAD is a public multimodal distraction dataset that VTD compares against, positioning VTD as more comprehensive in scenarios and modalities."},{"cited_title":"Dmd: A large-scale multi-modal driver monitoring dataset for attention and alertness analysis,","cited_arxiv_id":null,"evidence_quote":"DMD is a large-scale multi-modal driver monitoring dataset used as a comparison point for comprehensive fatigue and distraction data."},{"cited_title":"Fatigueview: A multi-camera video dataset for vision-based drowsiness detection,","cited_arxiv_id":null,"evidence_quote":"YawDD represents a single-modality yawning detection dataset that VTD aims to surpass with multi-modal synchronized data."},{"cited_title":"Real-time system for monitoring driver vigilance,","cited_arxiv_id":null,"evidence_quote":"FatigueView is a multi-camera drowsiness detection dataset cited to support the need for richer contextual data beyond single-modality visual features."}],"review_version":1}