Pith. sign in

REVIEW 4 major objections 5 minor 42 references

VTD: Visual and Tactile Database for Driver State and Behavior Perception

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read VTD is a new multimodal benchmark that synchronizes visual driver monitoring (frontal video, eye tracking) with tactile sensing (ECG from steering-wheel electrodes) and vehicle signals, built on 10 hours of induced-fatigue driving from 15…

desk verdict A promising dataset proposal whose central artifact—the data—is not actually available; the protocol detail is real, but the benchmark claim cannot be evaluated until VTD is released. read the letter →

arxiv 2412.04888 v1 pith:WFWWYN3V submitted 2024-12-06 cs.RO cs.AI

classification cs.ROcs.AI
keywords visualandtactiledatadriverstatebehaviorintelligentcockpitautonomousvehiclesfatiguetakeoverscenariosmultimodaldataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims to establish VTD, a multimodal dataset for studying driver fatigue and distraction in human-vehicle co-driving. It synchronizes frontal video, eye tracking, ECG acquired through steering-wheel electrodes, and vehicle signals, collected from 15 fatigued drivers (10 hours) and 17 distracted drivers (102 takeover trials). The central assertion is that this visual-tactile combination, with human-in-the-loop subjective labeling, forms a standardized benchmark for cross-modal driver state and behavior perception. A sympathetic reader would care because current public datasets are mostly single-modality and lack synchronized physiological and behavioral signals.

What carries the argument

The central mechanism is the dataset itself and its synchronized multi-modal acquisition platform: a driving simulator (Logitech G29) with flexible ECG electrodes built into the steering wheel, an RGB camera, Tobii Glasses3 eye tracking, and vehicle signal logging. The fatigue induction protocol defines four pre-fatigue states (A1 regular sleep, A2 nap deprivation, A3 partial night sleep deprivation, A4 total night sleep deprivation) and labels fatigue via KSS and SSS self-reports combined with staff observation. The takeover protocol defines three driving conditions and visual and auditory subtasks, measuring reaction time as the period from the takeover request to both hands returning to the wheel, and execution time from steering wheel and pedal thresholds.

What would settle it

If the released dataset's modality timestamps are not aligned within the stated frame rates (e.g., video at 30 fps versus ECG at 250 Hz with drift exceeding 100 ms), or if the KSS and SSS labels show no agreement with objective measures such as PERCLOS or heart-rate variability, the central claim of a synchronized, reliable benchmark would be undercut.

Watch

Extended reading notes

Core claim

On the paper's own terms, VTD is a new large-scale multimodal dataset that pairs visual signals (RGB facial video, eye tracking, head pose) with tactile and physiological signals (ECG via steering-wheel electrodes, vehicle control signals) under controlled fatigue and distraction protocols. It provides an 11-dimensional time series for fatigue, including EAR, PERCLOS, blinking rate, MAR, head tilt, R-R intervals, SDNN, RMSSD, steering angle, pedal, speed, and transverse angular velocity, plus takeover reaction and execution times. The paper argues that this fills the gap left by single-mode datasets and enables cross-modal fusion algorithms for driver monitoring.

Load-bearing premise

The fatigue labels come from KSS and SSS self-reports and staff observation rather than any objective physiological or behavioral ground truth, so the benchmark's value depends on those subjective ratings faithfully tracking true fatigue.

Editorial extensions

If this is right

  • Researchers can train cross-modal driver monitoring models that fuse facial video with ECG and vehicle signals using VTD's synchronized time series.
  • The 102 takeover scenarios with defined reaction and execution times allow benchmarking of takeover performance under visual versus auditory subtasks across different road geometries.
  • The steering-wheel ECG collection method offers a less intrusive alternative to wearable ECG sensors for driver monitoring, potentially transferable to real vehicles.
  • The 10 hours of fatigue data with 11-dimensional time series and a fixed 4:1:1 train-validation-test split support the development and comparison of fatigue detection algorithms.
  • VTD's multimodal design enables investigation of how visual features (e.g., PERCLOS, head tilt) and physiological features (e.g., SDNN, RMSSD) jointly indicate fatigue levels, going beyond single-modality approaches.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The steering-wheel ECG approach, if validated in real vehicles, could enable continuous driver monitoring without wearable sensors, but the simulator environment may not capture the motion artifacts and electrical noise of real-road driving.
  • Because fatigue labels rely on subjective KSS/SSS ratings and staff observation, benchmark users may need robust or semi-supervised methods to handle label noise, a direction the paper does not explore.
  • The combination of eye tracking and ECG could be extended beyond fatigue to estimate cognitive workload or engagement in non-driving tasks, which the paper mentions only through the QN-ACTR load-rate model.
  • Findings from this simulator-based dataset may transfer to real driving only after domain adaptation; the paper does not provide real-world validation, so users should treat simulator-to-real generalization as an open question.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents VTD, a multimodal dataset for driver state and behavior perception, collected in a driving simulator. For the fatigue portion, 15 participants drove for 40 minutes each under four pre-fatigue states while frontal RGB video, ECG from steering-wheel electrodes, eye tracking, and vehicle signals were recorded; labels were derived from KSS/SSS self-reports combined with staff observation. For the distraction/takeover portion, 17 participants performed 102 takeovers under visual and auditory subtasks across three road scenarios. The manuscript describes the data collection infrastructure, feature extraction (Mediapipe landmarks, EAR/MAR, head pose, heart-rate variability), a one-way ANOVA of time-series features against fatigue levels, and a comparison table against existing datasets. The intended contribution is that VTD is a comprehensive, synchronized, large-scale benchmark for cross-modal driver perception.

Significance. If the data were actually released and the labeling protocol validated, VTD would be a valuable addition to driver monitoring datasets, chiefly because the contactless steering-wheel ECG channel is combined with synchronized frontal video, eye tracking, and vehicle signals. The four-level pre-fatigue protocol and the repeated five-minute self/staff reporting cadence are a reasonable experimental design, and the paper usefully situates the work against existing DMS datasets. However, the benchmark claim cannot be assessed from the manuscript as it stands: the data are not accessible, no baseline experiments are run, the statistical evidence in Table III has internal inconsistencies, and the fatigue labels carry a circularity risk. The paper deserves a major revision rather than acceptance in its current form.

major comments (4)
  1. [IV.A and Abstract] The central claim that VTD is 'a comprehensive, well-structured, large-scale dataset' and a 'standardized platform for benchmarking' cannot be verified because the manuscript supplies no dataset release mechanism: there is no Data Availability statement, repository URL, DOI, download link, file manifest, sample data, or annotation example anywhere in Sections III–V. For a dataset paper, accessibility is the minimal condition for the benchmark claim; without it, the claimed 600 minutes of fatigue data and 102 takeover experiments remain a self-report. Please add a persistent public release (or a documented controlled-access procedure with an institutional contact), a file-format and folder-structure description, and at least one sample snippet of each modality, including the annotation format.
  2. [III.B and III.E] The fatigue labels are constructed from KSS/SSS self-reports and staff observation, and the same types of behavioral indicators (eye closure, head pose, facial action) are then tested in the Table III ANOVA. To the extent that the staff used those very signals to assign fatigue levels, the significant F statistics may reflect the labeling rule rather than an independent physiological correlate. The manuscript does not report inter-rater reliability, the number of label disagreements, the label distribution across the four pre-fatigue states, or any validation of the self-reports against objective performance. Please report these quantities and specify exactly how self-reports and staff observations were combined into the final labels.
  3. [Table III] The statistical evidence in Table III is internally inconsistent and therefore not load-bearing. The caption defines '++++' as α<0.01 and '+++' as 0.01≤α<0.05, but the Steering Angle(SD) row reports P=5.5569×10−2 ≈ 0.056, which is not significant at α=0.05 yet is marked '+++', while the Pedal(SD) row reports P=1.4349×10−2 and is marked '++++' even though it should be '+++' by the caption; the P value for Steering Angle(SD) also looks like a transcription of the Blinking Rate P value. In addition, the text alternates between '11-dimensional' (Section III.E) and '10-dimensional' (Table III, Table IV) time series. Please recompute the table, state the exact hypothesis and multiple-comparison correction, and confirm whether samples from the same driver are treated as independent in the ANOVA.
  4. [IV.A and Section III.E] The benchmark claim is unsupported by any experimental demonstration. Section III.E reports a 4:1:1 train/validation/test split and Table III gives univariate ANOVAs, but no classification or regression baseline is run on VTD, no metric (accuracy, F1, AUC) is reported, and no comparison with existing datasets such as DMD or FatigueView is made under a common protocol. To substantiate the 'benchmarking platform' contribution, add at least one well-defined baseline task (e.g., fatigue-level classification and takeover reaction-time prediction) with subject-independent evaluation and standard metrics.
minor comments (5)
  1. [Abstract and Table IV] The Abstract says '600 minutes' while Section IV.A says '10-hour fatigue driving data' and Table IV lists '630 min videos' for VTD; clarify whether the 630 minutes includes the takeover videos and whether the 600 minutes refers only to the fatigue protocol.
  2. [III.C] Section III.C states that '34 groups of visual and auditory subtasks' were established but Table IV reports '6 types of scenarios, 102 takeover experiments'; specify how the 34 groups and 6 scenario types relate to the three conditions (straight path, roundabout cut-in, roundabout obstacle avoidance).
  3. [III.B] The sentence 'The dataset labels were modified based on KSS and SSS fatigue scales' is vague; state which scale was collected at which time point and how the final ordinal label was computed.
  4. [III.B] The relation between 600 minutes of fatigue recording, 60-second time slices, and 480 valid samples is unexplained; specify the windowing, overlap, and exclusion criteria used to obtain 480 samples.
  5. [Equations (1)–(2)] The landmark indices P82, P87, P312, etc. are undefined; give a pointer to the canonical face model or include a landmark index diagram so that the MAR and EAR definitions are reproducible.

Circularity Check

1 steps flagged · score 3.0 of 10

The Table III ANOVA partly re-discovers the staff-observation cues used to assign fatigue labels; the dataset's descriptive claims are otherwise independent.

  1. self definitional [Section III.B (label collection), Section III.E (label grading/feature screening), Table III (ANOVA)]
    "the staff completed their assessments on the participants, combining the observation results and reported results ... Fatigue levels are then graded by subjective evaluations combined with self-assessment and other's assessment, thus realizing data calibration of Human-in-the-loop. ... The time series in some dimensions are chosen and investigated using One-way ANOVA ... to determine whether time series features are salient under different fatigue levels. From the ten features analyzed, VTD's fatigue data and driver fatigue are strongly correlated."

    The fatigue label is not an independent ground truth: it is assigned from KSS/SSS self-reports plus staff observation of the same frontal video, face, and driving behavior that are subsequently encoded as EAR, PERCLOS, blink rate, MAR, head pose, and vehicle signals. The ANOVA in Table III therefore partly tests features against a label that was itself informed by those features, so the reported 'strong correlation' is in part a reconstruction of the labeling signal rather than external validation. This does not invalidate the dataset's descriptive content, but it undercuts the evidential value of Table III as evidence that the features track fatigue.

full rationale

The paper's central contribution is a multimodal dataset, not a fitted model or derived prediction, so the main circularity failure modes do not apply. The only load-bearing circular element is the fatigue-label validation loop: labels are partly derived from staff observation of the same visual and behavioral cues that the ANOVA then reports as correlated with fatigue. The self-citation [41] used to support visual-tactile fusion is a normal literature reference and is not load-bearing for the dataset construction. Absence of a download link or repository is a verifiability/availability problem, not circularity. Overall: one partial, supporting circular step; the dataset's descriptive claims (modalities, synchronization, subject counts) remain independent.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central contribution is a dataset, so no fitted model parameters are introduced. The load-bearing assumptions are domain assumptions about label validity, simulator transferability, signal quality, and the face landmark pipeline. No new theoretical entities are postulated.

assumptions (5)
  • domain assumption Subjective KSS and SSS self-reports plus staff assessments provide valid fatigue labels.
    Section III.B labels all fatigue data using these subjective evaluations; no objective ground truth or inter-rater reliability is reported.
  • domain assumption ECG signals captured through steering-wheel electrodes during active driving are clean enough for R-R interval and HRV analysis after Butterworth filtering.
    Section III.B describes filtering but provides no signal-quality metrics, artifact rejection rates, or comparison with a reference ECG.
  • domain assumption Driving-simulator fatigue and distraction protocols induce states representative of real-world driving.
    Section III.C uses sparse simulated traffic and 40-minute sessions; transfer to on-road driving is asserted, not tested.
  • domain assumption Mediapipe Facemesh landmark positions can be mapped to MAR, EAR, and head-pose features under this camera setup.
    Section III.B depends on Mediapipe Facemesh and PnP pose estimation without reporting face-detection accuracy or failure cases.
  • standard math ANOVA F-statistics follow the F(k-1, n-k) distribution under the standard assumptions of the data.
    The footnote in Section III.E uses this standard statistical assumption to derive p-values for feature saliency.

how reviews work

0 comments
Cite this review

Pith. "Pith review of VTD: Visual and Tactile Database for Driver State and Behavior Perception." pith.science (2026). https://pith.science/paper/WFWWYN3V

@misc{pith2026241204888,
  author       = {Pith},
  title        = {Pith review of: VTD: Visual and Tactile Database for Driver State and Behavior Perception},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WFWWYN3V}},
  note         = {Machine review of arXiv:2412.04888}
}
read the original abstract

In the domain of autonomous vehicles, the human-vehicle co-pilot system has garnered significant research attention. To address the subjective uncertainties in driver state and interaction behaviors, which are pivotal to the safety of Human-in-the-loop co-driving systems, we introduce a novel visual-tactile perception method. Utilizing a driving simulation platform, a comprehensive dataset has been developed that encompasses multi-modal data under fatigue and distraction conditions. The experimental setup integrates driving simulation with signal acquisition, yielding 600 minutes of fatigue detection data from 15 subjects and 102 takeover experiments with 17 drivers. The dataset, synchronized across modalities, serves as a robust resource for advancing cross-modal driver behavior perception algorithms.

Figures

Figures reproduced from arXiv: 2412.04888 by the authors.

Figure 1
Figure 1. Research Gap in Driver Behavior and Driver State [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. VTD Data Collection Infrastructure the current status” every five minutes, the participants should complete their self-evaluations while the staff completed their assessments on the participants, combining the obser￾vation results and reported results. We generalized the different types of data collected as video data, vehicle data, and ECG data. The data were processed in different ways according to their character… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

42 extracted references · 39 canonical work pages

  1. [1]

    A simulation system for human-in-the-loop driving,

    Y . Li, Y . Su, X. Zhang, Q. Cai, H. Lu, and Y . Liu, “A simulation system for human-in-the-loop driving,” in 2022 IEEE 25th International Conference on Intelligent Transportation Systems (ITSC) . IEEE, 2022, pp. 4183–4188

  2. [2]

    Toward human-in-the-loop ai: Enhancing deep reinforcement learning via real-time human guidance for autonomous driving,

    J. Wu, Z. Huang, Z. Hu, and C. Lv, “Toward human-in-the-loop ai: Enhancing deep reinforcement learning via real-time human guidance for autonomous driving,” Engineering, vol. 21, pp. 75–91, 2023

  3. [3]

    Trial verification of human reliance on autonomous vehicles from the viewpoint of human factors,

    T. Arakawa, “Trial verification of human reliance on autonomous vehicles from the viewpoint of human factors,” Int. J. Innov. Comput. Inf. Control, vol. 14, no. January 2017, pp. 491–501, 2018

  4. [4]

    Analyzing driver behavior under natural- istic driving conditions: A review,

    H. Singh and A. Kathuria, “Analyzing driver behavior under natural- istic driving conditions: A review,” Accident Analysis & Prevention , vol. 150, p. 105908, 2021

  5. [5]

    Human-machine inter- action as key technology for driverless driving-a trajectory-based shared autonomy control approach,

    S. Gnatzig, F. Schuller, and M. Lienkamp, “Human-machine inter- action as key technology for driverless driving-a trajectory-based shared autonomy control approach,” in 2012 IEEE RO-MAN: The 21st IEEE International Symposium on Robot and Human Interactive Communication. IEEE, 2012, pp. 913–918

  6. [6]

    A review on driver face monitoring systems for fatigue and distraction detection,

    M.-H. Sigari, M.-R. Pourshahabi, M. Soryani, and M. Fathy, “A review on driver face monitoring systems for fatigue and distraction detection,” International Journal of Advanced Science and Technology, vol. 64, pp. 73–100, 2014

  7. [7]

    Head pose estimation in computer vision: A survey,

    E. Murphy-Chutorian and M. M. Trivedi, “Head pose estimation in computer vision: A survey,” IEEE transactions on pattern analysis and machine intelligence , vol. 31, no. 4, pp. 607–626, 2008

  8. [8]

    Ecg-based driver distraction iden- tification using wavelet packet transform and discriminative kernel- based features,

    S. V . Deshmukh and O. Dehzangi, “Ecg-based driver distraction iden- tification using wavelet packet transform and discriminative kernel- based features,” in 2017 IEEE International Conference on Smart Computing (SMARTCOMP) . IEEE, 2017, pp. 1–7

Show all 42 references
  1. [9]

    Ecg-based driver inattention identification during naturalistic driving using mel- frequency cepstrum 2-d transform and convolutional neural networks,

    M. Taherisadr, P. Asnani, S. Galster, and O. Dehzangi, “Ecg-based driver inattention identification during naturalistic driving using mel- frequency cepstrum 2-d transform and convolutional neural networks,” Smart health , vol. 9, pp. 50–61, 2018

  2. [10]

    A temporal–spatial deep learning approach for driver distraction detection based on eeg signals,

    G. Li, W. Yan, S. Li, X. Qu, W. Chu, and D. Cao, “A temporal–spatial deep learning approach for driver distraction detection based on eeg signals,” IEEE Transactions on Automation Science and Engineering , vol. 19, no. 4, pp. 2665–2677, 2021

  3. [11]

    Au- tomobile drivers distraction avoiding system using galvanic skin responses,

    P. Manikandan, M. R. S. Reddy, S. Mehatab, and P. M. Sai, “Au- tomobile drivers distraction avoiding system using galvanic skin responses,” in 2021 6th International Conference on Communication and Electronics Systems (ICCES) . IEEE, 2021, pp. 1818–1821

  4. [12]

    Driver distraction detection using bidirectional long short-term network based on multiscale entropy of eeg,

    X. Zuo, C. Zhang, F. Cong, J. Zhao, and T. H ¨am¨al¨ainen, “Driver distraction detection using bidirectional long short-term network based on multiscale entropy of eeg,” IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 10, pp. 19 309–19 322, 2022

  5. [13]

    Vision- based human action recognition: An overview and real world chal- lenges,

    I. Jegham, A. B. Khalifa, I. Alouani, and M. A. Mahjoub, “Vision- based human action recognition: An overview and real world chal- lenges,” F orensic Science International: Digital Investigation, vol. 32, p. 200901, 2020

  6. [14]

    Safe driving: Driver action recognition using surf keypoints,

    ——, “Safe driving: Driver action recognition using surf keypoints,” in 2018 30th International Conference on Microelectronics (ICM) . IEEE, 2018, pp. 60–63

  7. [15]

    A novel public dataset for multimodal multiview and multi- spectral driver distraction analysis: 3mdad,

    ——, “A novel public dataset for multimodal multiview and multi- spectral driver distraction analysis: 3mdad,” Signal Processing: Image Communication, vol. 88, p. 115960, 2020

  8. [16]

    Dmd: A large-scale multi-modal driver monitoring dataset for attention and alertness analysis,

    J. D. Ortega, N. Kose, P. Ca ˜nas, M.-A. Chao, A. Unnervik, M. Nieto, O. Otaegui, and L. Salgado, “Dmd: A large-scale multi-modal driver monitoring dataset for attention and alertness analysis,” in Computer Vision–ECCV 2020 Workshops: Glasgow, UK, August 23–28, 2020, Proceedin...

  9. [17]

    On performance evaluation of driver hand detection algorithms: Challenges, dataset, and metrics,

    N. Das, E. Ohn-Bar, and M. M. Trivedi, “On performance evaluation of driver hand detection algorithms: Challenges, dataset, and metrics,” in 2015 IEEE 18th international conference on intelligent transportation systems. IEEE, 2015, pp. 2953–2958

  10. [18]

    Driveahead-a large-scale driver head pose dataset,

    A. Schwarz, M. Haurilet, M. Martinez, and R. Stiefelhagen, “Driveahead-a large-scale driver head pose dataset,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2017, pp. 1–10

  11. [19]

    Driver anomaly detection: A dataset and contrastive learning approach,

    O. Kopuklu, J. Zheng, H. Xu, and G. Rigoll, “Driver anomaly detection: A dataset and contrastive learning approach,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 91–100

  12. [20]

    Mdad: A multimodal and multiview in-vehicle driver action dataset,

    I. Jegham, A. Ben Khalifa, I. Alouani, and M. A. Mahjoub, “Mdad: A multimodal and multiview in-vehicle driver action dataset,” in Com- puter Analysis of Images and Patterns: 18th International Conference, CAIP 2019, Salerno, Italy, September 3–5, 2019, Proceedings, Part I

  13. [21]

    Springer, 2019, pp. 518–529

  14. [22]

    Yawning detection using embedded smart cameras,

    M. Omidyeganeh, S. Shirmohammadi, S. Abtahi, A. Khurshid, M. Farhan, J. Scharcanski, B. Hariri, D. Laroche, and L. Martel, “Yawning detection using embedded smart cameras,” IEEE Transac- tions on Instrumentation and Measurement , vol. 65, no. 3, pp. 570– 582, 2016

  15. [23]

    The monitoring method of driver’s fatigue based on neural network,

    Y . Ying, S. Jing, and Z. Wei, “The monitoring method of driver’s fatigue based on neural network,” in 2007 International Conference on Mechatronics and Automation . IEEE, 2007, pp. 3555–3559

  16. [24]

    Real-time system for monitoring driver vigilance,

    L. M. Bergasa, J. Nuevo, M. A. Sotelo, R. Barea, and M. E. Lopez, “Real-time system for monitoring driver vigilance,”IEEE Transactions on intelligent transportation systems , vol. 7, no. 1, pp. 63–77, 2006

  17. [25]

    Fatigueview: A multi-camera video dataset for vision-based drowsiness detection,

    C. Yang, Z. Yang, W. Li, and J. See, “Fatigueview: A multi-camera video dataset for vision-based drowsiness detection,” IEEE Transac- tions on Intelligent Transportation Systems , vol. 24, no. 1, pp. 233– 246, 2022

  18. [26]

    Yawdd: A yawning detection dataset,

    S. Abtahi, M. Omidyeganeh, S. Shirmohammadi, and B. Hariri, “Yawdd: A yawning detection dataset,” in Proceedings of the 5th ACM multimedia systems conference , 2014, pp. 24–28

  19. [27]

    Eyeblink-based anti-spoofing in face recognition from a generic webcamera,

    G. Pan, L. Sun, Z. Wu, and S. Lao, “Eyeblink-based anti-spoofing in face recognition from a generic webcamera,” in 2007 IEEE 11th international conference on computer vision . IEEE, 2007, pp. 1–8

  20. [28]

    Driver drowsiness detection via a hierarchical temporal deep belief network,

    C.-H. Weng, Y .-H. Lai, and S.-H. Lai, “Driver drowsiness detection via a hierarchical temporal deep belief network,” in Computer Vision– ACCV 2016 Workshops: ACCV 2016 International Workshops, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part III 13 . Spring...

  21. [29]

    A realistic dataset and baseline temporal model for early drowsiness detection,

    R. Ghoddoosian, M. Galib, and V . Athitsos, “A realistic dataset and baseline temporal model for early drowsiness detection,” in Proceedings of the ieee/cvf conference on computer vision and pattern recognition workshops, 2019, pp. 0–0

  22. [30]

    A reduced feature set for driver head pose estimation,

    K. Diaz-Chito, A. Hern ´andez-Sabat´e, and A. M. L ´opez, “A reduced feature set for driver head pose estimation,” Applied Soft Computing , vol. 45, pp. 98–107, 2016

  23. [31]

    A review of utdrive studies: Learning driver behavior from naturalistic driving data,

    Y . Liu and J. H. Hansen, “A review of utdrive studies: Learning driver behavior from naturalistic driving data,” IEEE Open Journal of Intelligent Transportation Systems , vol. 2, pp. 338–346, 2021

  24. [32]

    Detecting human driver inattentive and aggressive driving behavior using deep learning: Recent advances, requirements and open challenges,

    M. H. Alkinani, W. Z. Khan, and Q. Arshad, “Detecting human driver inattentive and aggressive driving behavior using deep learning: Recent advances, requirements and open challenges,” Ieee Access , vol. 8, pp. 105 008–105 030, 2020

  25. [33]

    Modeling drowsy driving behaviors,

    X. Hu, R. Eberhart, and B. Foresman, “Modeling drowsy driving behaviors,” in Proceedings of 2010 IEEE International Conference on V ehicular Electronics and Safety . IEEE, 2010, pp. 13–17

  26. [34]

    Real-time facial surface geometry from monocular video on mobile gpus,

    Y . Kartynnik, A. Ablavatski, I. Grishchenko, and M. Grundmann, “Real-time facial surface geometry from monocular video on mobile gpus,” arXiv preprint arXiv:1907.06724 , 2019

  27. [35]

    Looking at faces in a vehicle: A deep cnn based approach and evaluation,

    K. Yuen, S. Martin, and M. M. Trivedi, “Looking at faces in a vehicle: A deep cnn based approach and evaluation,” in 2016 IEEE 19th International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2016, pp. 649–654

  28. [36]

    Vision- based method for detecting driver drowsiness and distraction in driver monitoring system,

    J. Jo, S. J. Lee, H. G. Jung, K. R. Park, and J. Kim, “Vision- based method for detecting driver drowsiness and distraction in driver monitoring system,” Optical Engineering, vol. 50, no. 12, pp. 127 202– 127 202, 2011

  29. [37]

    Modeling and recognition of driving fatigue state based on rr intervals of ecg data,

    L. Wang, J. Li, and Y . Wang, “Modeling and recognition of driving fatigue state based on rr intervals of ecg data,” Ieee Access, vol. 7, pp. 175 584–175 593, 2019

  30. [38]

    Fatigue driving detection model based on multi-feature fusion and semi-supervised active learn- ing,

    X. Li, L. Hong, J.-c. Wang, and X. Liu, “Fatigue driving detection model based on multi-feature fusion and semi-supervised active learn- ing,” IET Intelligent Transport Systems , vol. 13, no. 9, pp. 1401–1409, 2019

  31. [39]

    Driver state estimation by convolutional neural network using multimodal sensor data,

    S. Lim and J. H. Yang, “Driver state estimation by convolutional neural network using multimodal sensor data,” Electronics Letters , vol. 52, no. 17, pp. 1495–1497, 2016

  32. [40]

    Mobile-based wearable-type of driver fatigue detection by gsr and emg,

    L. Boon-Leng, L. Dae-Seok, and L. Boon-Giin, “Mobile-based wearable-type of driver fatigue detection by gsr and emg,” inTENCON 2015-2015 IEEE Region 10 Conference . IEEE, 2015, pp. 1–4

  33. [41]

    Wearable glove-type driver stress detection using a motion sensor,

    B.-G. Lee and W.-Y . Chung, “Wearable glove-type driver stress detection using a motion sensor,” IEEE Transactions on Intelligent Transportation Systems, vol. 18, no. 7, pp. 1835–1844, 2016

  34. [42]

    A survey of driver behavior perception methods for human-computer hybrid enhancement of intel- ligent driving,

    J. Yi, A. Du, Z. Zhu, and H. Ding, “A survey of driver behavior perception methods for human-computer hybrid enhancement of intel- ligent driving,” in Proceedings of China SAE Congress 2021: Selected Papers. Springer, 2022, pp. 754–766

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.