REVIEW 4 major objections 4 minor 55 references
Driver Behavior Under Traffic Complexity: Variance Decomposition and Cross-Driver Generalization for Driver Monitoring
T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Traffic complexity leaves only a 1.5% trace in driver behavior, while driver identity accounts for 23% of the variance—a 15:1 imbalance that caps all tested models at F1≈0.45 for unseen drivers.
desk verdict Thorough feature screen and an honest negative result, but the headline 15:1 variance ratio leans on expert labels whose validity the authors themselves hedge; referee it, don't treat it as a design constraint yet. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a linear mixed-effects variance decomposition (random intercept and random slope per driver) applied to the 31 surviving features. It partitions each feature's variance into a complexity component (marginal R²), a driver-identity component (intraclass correlation), and a driver-specific sensitivity component (slope share). The paper shows these three numbers predict downstream performance: the 15:1 identity-to-complexity ratio matches the cross-driver generalization gap and the uniform F1 ceiling. A secondary mechanism is the two-stage screening pipeline—non-parametric tests with false-discovery-rate correction, speed residualization, equivalence testing, and dr
What would settle it
Collect a dataset where drivers rate complexity continuously in situ (or labels come from a validated objective benchmark) and run the same mixed-effects decomposition. If complexity's variance share rises substantially above ~1.5%, or if any classifier breaks clearly above F1≈0.45 under leave-one-driver-out on that data, the variance-structure bottleneck claim fails. A cheaper check within the same data: test whether the 1.5% share is reproduced using only labels from the single rater with highest agreement with the consensus, and whether it shrinks when the 'needed adjustments' phrasing is r
Extended reading notes
Core claim
The central discovery is a variance-structure bottleneck: for behavioral metrics that genuinely respond to traffic complexity, the proportion of variance attributable to complexity is tiny (median marginal R² ≈ 0.015) compared with stable driver identity (median ICC ≈ 0.228), leaving roughly 75% residual variance. The paper interprets the resulting 15:1 ratio as the common cause of the failure of all eight feature-level personalization strategies (even oracle normalization cannot recover a signal that is mostly residual) and of the convergence of four classification architectures (gradient boosting, SVM, GRU with attention, logistic regression) at macro F1 ≈ 0.45 under leave-one-subject-out
Load-bearing premise
The expert-labeled traffic-complexity ground truth (three ordinal classes rated from front-camera video, with only moderate inter-rater agreement, α=0.57) is a valid, behavior-independent measure of situational demand; if the labels already encode the behaviors being tested, part of the reported 1.5% complexity share could be an artifact rather than a true signal.
Editorial extensions
If this is right
- Complexity-adaptive ADAS cannot rely on behavioral features alone for a previously unseen driver; population-level models cap at F1≈0.45, above the majority-class baseline but below standalone deployment.
- Feature-level personalization (z-score or mean-subtraction, any temporal scope) is ruled out as a recovery path; model-level calibration is the required alternative.
- Guiding fixation rate, combined with braking-entropy features, offers the best driver-agnostic signal and should be prioritized in driver-monitoring feature sets.
- Deployment splits into three regimes: driver-agnostic (universal features only), personalized (few-shot calibration), and hybrid (behavior plus external context such as route priors or traffic density).
- Because residual variance dominates at ~75%, even perfect removal of driver identity leaves limited recoverable complexity signal; the bottleneck is structural, not architectural.
Reading between the lines
- If the 15:1 ratio holds more broadly, then estimating mental workload or task demand from driver behavior in real vehicles faces the same ceiling, and fusing behavioral cues with environmental measures becomes a necessity rather than an enhancement.
- The observed yaw-pitch asymmetry and the dominance of entropy over mean-level features suggest complexity acts on the temporal regularity of attention rather than its spatial average; a testable extension would be to check whether the same asymmetry appears in non-driving vigilance tasks.
- The slope-share analysis implies that a small per-driver calibration set (e.g., a few minutes) should pay off most for braking-entropy features and least for guiding fixation rate; this ordering is directly testable with a few-shot meta-learning baseline.
- Because labels were assigned post hoc from video, an in-situ probe (e.g., continuous subjective difficulty ratings collected during the drive) could either confirm the 1.5% figure or reveal that part of the residual variance is label noise; the paper itself concedes this overlap cannot be ruled out.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether traffic complexity can be inferred from driver behavior in partially automated urban driving. Using data from 20 drivers, the authors screen 175 behavioral features across five domains with a two-stage pipeline (sample-level Kruskal-Wallis with FDR and TOST equivalence testing, then driver-level tier ranking). They retain 31 features and fit linear mixed-effects models to decompose variance into complexity-related, driver-identity, and residual components, reporting a median R²_m of 0.015 for complexity versus a median ICC of 0.228 for driver identity—a roughly 15:1 ratio. The paper then evaluates eight feature-level personalization strategies (all of which degrade discrimination) and benchmarks four classifiers under leave-one-subject-out and block-aware cross-validation, finding that all architectures converge at macro-F1 ≈ 0.45 under LOSO. The authors interpret this as evidence that variance structure, not model capacity, bounds complexity estimation, and they propose guiding fixation rate as the most deployment-ready feature.
Significance. If the central variance-structure claim holds, the paper makes a substantial contribution: it is one of the first systematic multi-domain screens of driver behavior for complexity estimation, it uses production-grade DMS sensors, and it provides a transparent, falsifiable account of why cross-driver complexity inference is hard. The double-gated screening, formal equivalence testing of null features, standard LOSO protocol, and explicit reporting of mixed-effects parameters are all strengths. The analysis is also unusually honest about its own limitations, including the moderate expert-label reliability and the admitted criterion overlap. However, the central quantitative claims—the 15:1 variance ratio and the F1≈0.45 ceiling—rest entirely on the validity of the expert complexity labels as a behavior-independent ground truth, and the paper's own acknowledgments leave this validity unresolved. The manuscript is therefore a strong candidate for major revision: the methodology is sound, but the load-bearing assumptions need additional analysis before the headline conclusions can be accepted.
major comments (4)
- [Section V-F and Section III-A] The central claim—that complexity explains only 1.5% of behavioral variance while driver identity explains 23%, yielding a 15:1 ratio and bounding cross-driver F1—depends on the expert labels being a valid, behavior-independent measure of traffic complexity. The authors themselves state in Section V-F that 'criterion overlap with the analyzed behavioral metrics cannot be ruled out', and the label definition in Section III-A includes 'needed adjustments' and 'frequent adjustments', which are behavioral concepts. This is not a peripheral caveat; it is load-bearing. If labels encode the same adjustments used as features, the small R²_m for complexity could be an artifact of label-behavior overlap, or conversely, label noise (Krippendorff's α=0.57) could attenuate the true complexity signal. I request a concrete re-analysis that either (a) uses a label protocol restricted to scene-based crit
- [Table I and Section III-D] Three retained features—BR_t (ICC=1.161), SD(Ψ_g) (ICC=1.340), and SD(Ψ_h) (ICC=1.376)—have ICC values exceeding unity. The paper attributes this to strongly negative intercept-slope covariance in the random-effects model (Eq. 2). However, an ICC outside [0,1] indicates that the variance partition is not identified for those features, and this is not merely a cosmetic issue: BR_t and SD(Ψ_g) are among the top SHAP features, and the median ICC of 0.228 and the 15:1 ratio depend on the entire set of retained features. I ask the authors to report the estimated intercept-slope covariance for all features, to show the distribution of the denominator of the ICC, and to provide a sensitivity analysis in which these ill-identified features are either excluded or reparameterized (e.g., constrained covariance or random-intercept-only models). The current text mentions a random-intercept-only sensi
- [Section IV-D and Section III-F] The authors interpret the convergence of all four architectures at macro-F1 ≈ 0.45 under LOSO as evidence that the variance structure, rather than model capacity, bounds complexity estimation. An alternative and equally parsimonious explanation is that the expert labels themselves are noisy (α=0.57), and any classifier trained on these labels is capped by label reliability. The two hypotheses—variance-structure ceiling versus label-noise ceiling—have different implications for the deployment regimes proposed in Section V-E. I request a concrete analysis that separates them: for example, train a classifier on one rater's labels and evaluate against another rater's labels (or against the Dawid-Skene aggregate), and compare the resulting F1 with the observed LOSO F1. If the label-noise ceiling is already near 0.45, the claim that the ceiling is imposed by the variance structure is not suppo
- [Section III-F and Section IV-D] The η²-based feature screening is performed on the full dataset, not nested inside the cross-validation loop. The authors argue the risk of overfitting is minimal because the screen is a univariate population-level filter with a theory-driven threshold. However, under LOSO the held-out driver's data are included in the feature-selection step, which can inflate the apparent cross-driver generalizability of the retained feature set and thus bias the LOSO F1 upward. This directly affects the interpretation of the F1≈0.45 ceiling as a bound imposed by variance structure. I ask for a robustness check in which the screening is repeated inside each LOSO fold, using only the 19 training drivers, with the held-out driver's features selected according to the training-only screen. If the F1 ceiling is unchanged, the conclusion is strengthened; if it degrades, the reported ceiling is partly an artif
minor comments (4)
- [Section III-D] The term 'slope share' is defined as σ²_u1 / (σ²_u0 + σ²_u1), but the variance decomposition framework of Nakagawa and Schielzeth uses a total variance that includes the intercept-slope covariance. This inconsistency should be stated explicitly when interpreting slope shares; otherwise readers may assume the quantity is a proper proportion bounded by [0,1].
- [Table I] The table reports 'Rob.' as ✓/✗ but the caption defines these as speed-robust/speed-mediated. Consider adding a footnote that clarifies that ✓ means the effect survives speed residualization, since the term 'robust' might be confused with statistical robustness.
- [Section V-C] The discussion of personalization says 'for universal features, no large between-driver offset exists to normalize.' This is slightly misleading given that universal features are defined by ICC<0.5, not by zero between-driver variance. A universal feature can still have substantial ICC below 0.5; the statement should be qualified.
- [Section IV-A] The sentence 'Features with η²_driver ≥0.01 but p_driver ≥0.05 as Tier 2 (power-limited, retained because the effect is present but statistical power is insufficient at N=60)' is a post-hoc power argument. It would be more precise to state that Tier 2 features are retained despite not reaching significance at the driver level, and to acknowledge the risk of false positives among them.
Circularity Check
Central variance decomposition is an honest measurement, but the expert-label ground truth is partly defined by the behaviors it is compared against; the authors concede the overlap, making the 15:1 ratio partially dependent on label definition.
-
self definitional
[Section III-A (labeling) and Section V-F (Limitations)]
"Three experts ... labeled each drive using three ordinal complexity classes: simple (clearly structured, low attention demand), normal (partially structured, minor adjustments needed), and complex (multifaceted, frequent adjustments, heightened attention). ... Because the label definition already included statements about needed adjustments, criterion overlap with the analyzed behavioral metrics cannot be ruled out."
The complexity ground truth used throughout the paper — for screening, for the mixed-effects variance decomposition (Eq. 2 via Nakagawa R2m), and for all classification labels — is operationalized with the phrases 'minor adjustments needed' and 'frequent adjustments.' The very behavioral metrics that the paper analyzes (braking intensity, speed-limit deviation, and other longitudinal-control adjustment features) are also 'adjustments.' Thus a portion of the reported 1.5% complexity-explained variance and the LOSO F1 ceiling may be an artifact of the label definition encoding the outcome behaviors, not an independent exogenous complexity signal. The authors explicitly concede this overlap, so the central ratio is not fully independent of its input definition.
full rationale
The paper's derivation chain is largely empirical rather than constructed: features are screened, variance is decomposed with mixed-effects models, personalization strategies are evaluated on the same features, and classifiers are benchmarked under honest LOSO cross-validation. I find no fitted parameter being renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no equation that reduces to its own input. The self-citations to the authors' preceding conference paper [13] are descriptive (dataset, prior labeling protocol, earlier 16-metric analysis) and are not used to forbid alternatives or to supply a derivational premise, so they do not constitute load-bearing circularity. The one substantive circularity risk is the expert-label ground truth: the complexity class definitions include 'minor adjustments needed' and 'frequent adjustments,' which are conceptually the same constructs as the analyzed braking and speed-adjustment behaviors. The paper itself states in Section V-F that 'criterion overlap with the analyzed behavioral metrics cannot be ruled out.' This makes a portion of the 1.5% complexity share, and hence the 15:1 driver-identity-to-complexity ratio, potentially definitional rather than a measurement of an exogenous complexity signal. Because this is an openly disclosed construct-validity limitation and not a deliberate construction or a fitted-prediction trick, the appropriate score is 4 rather than 6 or higher. The core statistical analyses would be self-contained and informative if the labels were validated as behavior-independent; the admitted overlap is the main reason the central claim does not receive a clean 0-2 score.
Assumptions & free parameters
free parameters (5)
- η² screening threshold =
0.01
- TOST equivalence bound Δ =
0.2 pooled SD (Cohen's d)
- Label alignment tolerance =
±2 s
- Complexity ordinal coding =
x = 0, 1, 2
- Rolling window length =
3 s at 10 Hz
assumptions (5)
- domain assumption Expert Dawid-Skene aggregated labels are a valid proxy for subjective traffic complexity.
- domain assumption Labels are independent of the behavioral features under test.
- domain assumption Linear mixed-effects model with Gaussian random intercept/slope is correctly specified for variance decomposition.
- domain assumption Driver-level aggregation to per-driver per-class medians (N=60) sufficiently corrects temporal autocorrelation in the Kruskal-Wallis tests.
- domain assumption Speed residualization via OLS removes the speed confound without removing complexity-driven behavior.
Cite this review
Pith. "Pith review of Driver Behavior Under Traffic Complexity: Variance Decomposition and Cross-Driver Generalization for Driver Monitoring." pith.science (2026). https://pith.science/paper/OTQFEP66
@misc{pith2026260716847,
author = {Pith},
title = {Pith review of: Driver Behavior Under Traffic Complexity: Variance Decomposition and Cross-Driver Generalization for Driver Monitoring},
year = {2026},
howpublished = {\url{https://pith.science/paper/OTQFEP66}},
note = {Machine review of arXiv:2607.16847}
}
read the original abstract
Understanding why a traffic situation is demanding for the human driver is central to safe and comfortable partially automated driving. Existing complexity metrics characterize only external environmental factors, while driver monitoring systems detect only endogenous states such as distraction. Neither side captures how external demands translate into driver-perceived situational load. Driver behavior serves as the natural bridge between both, and this paper investigates the inverse inference problem of estimating traffic complexity from behavioral signals. A systematic screening of 175 behavioral features across five domains (gaze, head pose, longitudinal control, guiding fixation, scanning strategy) is applied to data from 20 drivers in real urban traffic. Of these, 31 features exhibit statistically confirmed complexity effects, while 140 are confirmed as null effects through equivalence testing. Mixed-effects variance decomposition reveals that complexity explains only 1.5% of behavioral variance, whereas driver identity accounts for 23% and residual variance for 75%. This unfavorable ratio explains both the failure of all eight evaluated feature-level personalization strategies and the convergence of four classification architectures at F1 around 0.45 under leave-one-subject-out cross-validation. Guiding fixation rate emerges as the single most deployment-ready feature, combining speed-robustness, universality across drivers, and minimal inter-driver variation in complexity sensitivity. The results define three deployment regimes for complexity-adaptive advanced driver assistance systems and establish the variance structure as the primary bottleneck for complexity estimation.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
The challenges of next-gen adas and ads and re- lated vehicle safety topics,
S. Chalmers, “The challenges of next-gen adas and ads and re- lated vehicle safety topics,” SAE International, SAE Research Report EPR2025003, 2025
2025
-
[2]
SAE J3016: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,
SAE International, “SAE J3016: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” p. 41, 2021
2021
-
[3]
UN Regulation No. 171: UN Regulation on uniform provisions concerning the approval of vehicles with regard to Driver Control Assistance Systems,
United Nations Economic Commission for Europe, “UN Regulation No. 171: UN Regulation on uniform provisions concerning the approval of vehicles with regard to Driver Control Assistance Systems,” 2024
2024
-
[4]
Safe driving driver engagement,
Euro NCAP, “Safe driving driver engagement,” European New Car Assessment Program, Protocol v1.1, 2025
2025
-
[5]
Auditory interfaces in automated driving: an international survey,
P. Bazilinskyy and J. de Winter, “Auditory interfaces in automated driving: an international survey,”PeerJ Comput. Sci., vol. 1:e13, pp. 1–28, 2015
2015
-
[6]
Analysis of Freeway Safety Influencing Factors on Driving Workload and Performance Based on the Gray Correlation Method,
L. Xie, C. Wu, M. Duan, and N. Lyu, “Analysis of Freeway Safety Influencing Factors on Driving Workload and Performance Based on the Gray Correlation Method,”J. Adv. Transp., vol. 2021, p. 1–11, 2021
2021
-
[7]
Driver’s Cognitive Workload and Driving Performance under Traffic Sign Information Exposure in Complex Environments: A Case Study of the Highways in China,
N. Lyu, L. Xie, C. Wu, Q. Fu, and C. Deng, “Driver’s Cognitive Workload and Driving Performance under Traffic Sign Information Exposure in Complex Environments: A Case Study of the Highways in China,”Int. J. Environ. Res. Public Health, vol. 14, p. 203, 2017
2017
-
[8]
Complexity Quantification of Driving Scenarios with Dynamic Evolution Character- istics,
T. Liu, C. Wang, Z. Yin, Z. Mi, X. Xiong, and B. Guo, “Complexity Quantification of Driving Scenarios with Dynamic Evolution Character- istics,”Entropy, vol. 26, no. 12, 2024
2024
Show all 55 references
-
[9]
Attention-based Neural Network for Driving Environment Complexity Perception,
C. Zhang, A. Eskandarian, and X. Du, “Attention-based Neural Network for Driving Environment Complexity Perception,” in2021 IEEE Inter- national Intelligent Transportation Systems Conference (ITSC), 2021, p. 2781–2787
2021
-
[10]
The characteristics of driver lane-changing behaviour in congested road environments,
W. Wang and G. Cheng, “The characteristics of driver lane-changing behaviour in congested road environments,”Transp. Saf. Environ., vol. 6, pp. 1–15, 2023
2023
-
[11]
Gaze- based indicators of driver cognitive distraction: Effects of different traffic conditions and adaptive cruise control use,
A. Halin, A. Deli `ege, C. Devue, and M. Van Droogenbroeck, “Gaze- based indicators of driver cognitive distraction: Effects of different traffic conditions and adaptive cruise control use,” inAdjunct Proceedings of the 17th International Conference on Automotive User Interfac...
2025
-
[12]
Analyzing Gaze During Driving: Should Eye Tracking Be Used to Design Automotive Lighting Functions?
K. Kunst, D. Hoffmann, A. Erkan, K. Lazarova, and T. Q. Khanh, “Analyzing Gaze During Driving: Should Eye Tracking Be Used to Design Automotive Lighting Functions?”J. Eye Mov. Res., vol. 18, p. 13, 2025
2025
-
[13]
Investigating Driver Be- havior in Complex Traffic Situations While Driving Partially Automated Vehicles,
L. K ¨oning, N. Mili ˇci´c, and K. Bogenberger, “Investigating Driver Be- havior in Complex Traffic Situations While Driving Partially Automated Vehicles,” 2026, in press, arXiv, 10.48550/arXiv.2607.00855
-
[14]
From Concept to Representation: Modeling Driving Capability and Task Demand with a Multimodal Large Language Model,
H. Zhou, A. Carballo, K. Fujii, and K. Takeda, “From Concept to Representation: Modeling Driving Capability and Task Demand with a Multimodal Large Language Model,”Sensors, vol. 25, no. 18, p. 5805, 2025
2025
-
[15]
Familiarity and Com- plexity during a Takeover in Highly Automated Driving,
M. S. L. Scharfe-Scherf and N. Russwinkel, “Familiarity and Com- plexity during a Takeover in Highly Automated Driving,”Int. J. Intell. Transp. Syst. Res., vol. 19, p. 525–538, 2021
2021
-
[16]
Research on the Quantitative Eval- uation of the Traffic Environment Complexity for Unmanned Vehicles in Urban Roads,
S. Yang, L. Gao, Y . Zhao, and X. Li, “Research on the Quantitative Eval- uation of the Traffic Environment Complexity for Unmanned Vehicles in Urban Roads,”IEEE Access, vol. 9, p. 23139–23152, 2021
2021
-
[17]
The task-capability interface model of the driving process,
R. Fuller, “The task-capability interface model of the driving process,” Rech. - Transp. - Secur., vol. 66, p. 47–57, 2000
2000
-
[18]
Mental workload and driving,
J. Paxion, E. Galy, and C. Berthelon, “Mental workload and driving,” Front. Psychol., vol. 5, p. 1–11, 2014
2014
-
[19]
Traffic Risk Envi- ronment Impact Analysis and Complexity Assessment of Autonomous Vehicles Based on the Potential Field Method,
Y . Cheng, Z. Liu, L. Gao, Y . Zhao, and T. Gao, “Traffic Risk Envi- ronment Impact Analysis and Complexity Assessment of Autonomous Vehicles Based on the Potential Field Method,”Int. J. Environ. Res. Public Health, vol. 19, no. 16, p. 10337, 2022
2022
-
[20]
Criticality Metrics for Automated Driving: A Review and Suitability Analysis of the State of the Art,
L. Westhofen, C. Neurohr, T. Koopmann, M. Butz, B. Sch ¨utt, F. Utesch, B. Neurohr, C. Gutenkunst, and E. B ¨ode, “Criticality Metrics for Automated Driving: A Review and Suitability Analysis of the State of the Art,”Arch. Comput. Methods Eng., vol. 30, no. 1, p. 1–35, 2023
2023
-
[21]
Objective-Subjective Consistent Evaluation of Scene Complexity Based on Multiscale Fuzzy Entropy,
H. Cao, Q. Meng, B. Li, and L. Zhang, “Objective-Subjective Consistent Evaluation of Scene Complexity Based on Multiscale Fuzzy Entropy,” in2025 9th CAA International Conference on Vehicular Control and Intelligence (CVCI), 2025, p. 1–6. 11
2025
-
[22]
Temporal fluctuations in driving demand: The effect of traffic complexity on subjective measures of workload and driving performance,
E. Teh, S. Jamson, O. Carsten, and H. Jamson, “Temporal fluctuations in driving demand: The effect of traffic complexity on subjective measures of workload and driving performance,”Transp. Res. Part F Traffic Psychol. Behav., vol. 22, p. 207–217, 2014
2014
-
[23]
Determining Infrastructure- and Traffic Factors that Increase the Perceived Complexity of Driving Situations,
A. Boelhouwer, A. P. van den Beukel, M. C. van der V oort, and M. H. Martens, “Determining Infrastructure- and Traffic Factors that Increase the Perceived Complexity of Driving Situations,” inAdvances in Human Aspects of Transportation. Springer, 2020, p. 3–10
2020
-
[24]
Driving task analysis as a tool in traffic safety research and practice,
W. Fastenmeier and H. Gstalter, “Driving task analysis as a tool in traffic safety research and practice,”Saf. Sci., vol. 45, no. 9, p. 952–979, 2007
2007
-
[25]
Looking at Humans in the Age of Self-Driving and Highly Automated Vehicles,
E. Ohn-Bar and M. M. Trivedi, “Looking at Humans in the Age of Self-Driving and Highly Automated Vehicles,”IEEE Trans. Intell. Veh., vol. 1, no. 1, p. 90–104, 2016
2016
-
[26]
Technologies for detecting and monitoring drivers’ states: A systematic review,
M. S. Al-Quraishi, S. S. Azhar Ali, M. Al-Qurishi, T. B. Tang, and S. Elferik, “Technologies for detecting and monitoring drivers’ states: A systematic review,”Heliyon, vol. 10, no. 20, p. e39592, 2024
2024
-
[27]
Head and eye gaze dynamics during visual attention shifts in complex environments,
A. Doshi and M. M. Trivedi, “Head and eye gaze dynamics during visual attention shifts in complex environments,”J. Vis., vol. 12, no. 2, 2012
2012
-
[28]
A review of gaze entropy as a measure of visual scanning efficiency,
B. Shiferaw, L. Downey, and D. Crewther, “A review of gaze entropy as a measure of visual scanning efficiency,”Neurosci. Biobehav. Rev., vol. 96, p. 353–366, 2019
2019
-
[29]
Gaze entropy metrics for mental workload estimation are heterogenous during hands- off level 2 automation,
C. M. Goodridge, R. C. Gonc ¸alves, A. Arabian, A. Horrobin, A. Sol- ernou, Y . T. Lee, Y . M. Lee, R. Madigan, and N. Merat, “Gaze entropy metrics for mental workload estimation are heterogenous during hands- off level 2 automation,”Accid. Anal. Prev., vol. 202, pp. 1–13, 2024
2024
-
[30]
Discriminative Capabilities of Eye Gaze Measures for Cognitive Load Evaluation in a Driving Simulation Task,
A. Bakhchina, K. Arutyunova, E. Burashnikov, A. Filatova, A. Fil- imonov, and I. Shishalov, “Discriminative Capabilities of Eye Gaze Measures for Cognitive Load Evaluation in a Driving Simulation Task,” J. Eye Mov. Res., vol. 19, no. 1, 2025
2025
-
[31]
Eye Gaze Entropy Reflects Individual Experience in the Context of Driving,
K. Arutyunova, E. Burashnikov, N. Timakin, I. Shishalov, A. Filimonov, and A. Bakhchina, “Eye Gaze Entropy Reflects Individual Experience in the Context of Driving,”Entropy, vol. 28, no. 1, 2025
2025
-
[32]
Driver Visual Attention Estimation Using Head Pose and Eye Appearance Information,
S. Jha, N. Al-Dhahir, and C. Busso, “Driver Visual Attention Estimation Using Head Pose and Eye Appearance Information,”IEEE Open J. Intell. Transp. Syst., vol. 4, p. 216–231, 2023
2023
-
[33]
Eye-head coordination and dynamic visual scanning as indicators of visuo-cognitive demands in driving simulator,
L. Mikula, S. Mej ´ıa-Romero, R. Chaumillon, A. Patoine, E. Lugo, D. Bernardin, and J. Faubert, “Eye-head coordination and dynamic visual scanning as indicators of visuo-cognitive demands in driving simulator,”PLoS One, vol. 15, no. 12, p. e0240201, 2020
2020
-
[34]
Where we look when we steer,
M. F. Land and D. N. Lee, “Where we look when we steer,”Nature, vol. 369, p. 742–4, 1994
1994
-
[35]
Gaze Strategies in Driving–An Ecological Approach,
O. Lappi, “Gaze Strategies in Driving–An Ecological Approach,”Front. Psychol., vol. 13, pp. 1–15, 2022
2022
-
[36]
Drivers use active gaze to monitor waypoints during automated driving,
C. Mole, J. Pekkanen, W. E. A. Sheppard, G. Markkula, and R. M. Wilkie, “Drivers use active gaze to monitor waypoints during automated driving,”Sci. Rep., vol. 11, pp. 1–18, 2021
2021
-
[37]
Behavioral Entropy as a Measure of Driving Performance,
E. Boer, “Behavioral Entropy as a Measure of Driving Performance,” in Driving Assessment Conference Archive, vol. 1, 2001, p. 226–229
2001
-
[38]
Towards recognizing cognitive distraction levels with low-cost and high-sensitive measures: The effectiveness of sample, approximate, and traditional steering entropies,
N. Chang, C. Dong, S. Zhu, P. Li, X. Li, L. Xia, and X. Yan, “Towards recognizing cognitive distraction levels with low-cost and high-sensitive measures: The effectiveness of sample, approximate, and traditional steering entropies,”Transp. Res. Part F Traffic Psychol. Behav., ...
2025
-
[39]
Domain Adaptive Driver Distraction Detection Based on Partial Feature Alignment and Confusion-Minimized Classification,
G. Li, G. Wang, Z. Guo, Q. Liu, X. Luo, B. Yuan, M. Li, and L. Yang, “Domain Adaptive Driver Distraction Detection Based on Partial Feature Alignment and Confusion-Minimized Classification,”IEEE Trans. Intell. Transp. Syst., vol. 25, no. 9, p. 11227–11240, 2024
2024
-
[40]
Beyond “One-Size-Fits-All
J. C. Pe ˜na, E. V ´asquez, G. A. Feo-Cediel, A. Negroni, and J. F. Medina-Lee, “Beyond “One-Size-Fits-All”: Estimating Driver Attention with Physiological Clustering and LSTM Models,”Electronics, vol. 14, no. 23, p. 4655, 2025
2025
-
[41]
When Is It Safe to Complete an Overtaking Maneuver? Modeling Drivers’ Decision to Return After Passing a Cyclist,
A. Rasch, C. Flannagan, and M. Dozza, “When Is It Safe to Complete an Overtaking Maneuver? Modeling Drivers’ Decision to Return After Passing a Cyclist,”IEEE Trans. Intell. Transp. Syst., vol. 25, no. 11, p. 15587–15599, 2024
2024
-
[42]
Correcting Time-Continuous Emotional Labels by Modeling the Reaction Lag of Evaluators,
S. Mariooryad and C. Busso, “Correcting Time-Continuous Emotional Labels by Modeling the Reaction Lag of Evaluators,”IEEE Trans. Affect. Comput., vol. 6, p. 97–108, 2015
2015
-
[43]
Eye tracking cognitive load using pupil diameter and microsaccades with fixed gaze,
K. Krejtz, A. T. Duchowski, A. Niedzielska, C. Biele, and I. Krejtz, “Eye tracking cognitive load using pupil diameter and microsaccades with fixed gaze,”PLoS One, vol. 13, no. 9, p. e0203629, 2018
2018
-
[44]
Driver Personalized Lane Change Behavior Analysis Based on Attention Deep Embedding Clustering,
H. Dong, W. Wang, Y . Wang, L. Li, Y . Yue, J. Tian, and J. Han, “Driver Personalized Lane Change Behavior Analysis Based on Attention Deep Embedding Clustering,” inSAE 2025 Intelligent and Connected Vehicles Symposium. SAE International, 2025
2025
-
[45]
Few-shot driver identification via meta-learning,
L. Lu and S. Xiong, “Few-shot driver identification via meta-learning,” Expert Syst. Appl., vol. 203, p. 117299, 2022
2022
-
[46]
A unified approach to interpreting model predictions,
S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” inProceedings of the 31st International Conference on Neural Information Processing Systems. Curran Associates Inc., 2017, p. 4768–4777
2017
-
[47]
Driver engagement and fully shared longitudinal control - insights from real partially automated vehicles,
L. K ¨oning, P. Santos, N. Milicic, and K. Bogenberger, “Driver engagement and fully shared longitudinal control - insights from real partially automated vehicles,” 2026, techRxiv, 10.36227/techrxiv.177126589.99792150/v1
2026
-
[48]
Improving Driver Engagement with Level 2 Automated Systems: The Impact of Fully Shared Longitudinal Control,
J. Illgner, N. Milicic, B. Biebl, and M. Baumann, “Improving Driver Engagement with Level 2 Automated Systems: The Impact of Fully Shared Longitudinal Control,” inProceedings of the 16th International Conference on Automotive User Interfaces and Interactive Vehicular Applicati...
2024
-
[49]
Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm,
A. P. Dawid and A. M. Skene, “Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm,”J. R. Stat. Soc. Ser. C Appl. Stat., vol. 28, p. 20–28, 1979
1979
-
[50]
The measurement of observer agreement for categorical data,
J. R. Landis and G. G. Koch, “The measurement of observer agreement for categorical data,”Biometrics, vol. 33, p. 159–74, 1977
1977
-
[51]
Cohen,Statistical Power Analysis for the Behavioral Sciences
J. Cohen,Statistical Power Analysis for the Behavioral Sciences. New York, NY , USA: Routledge, 1988, vol. 2
1988
-
[52]
Controlling the false discovery rate: A practical and powerful approach to multiple testing,
Y . Benjamini and Y . Hochberg, “Controlling the false discovery rate: A practical and powerful approach to multiple testing,”J. R. Stat. Soc. Ser. B Methodol., vol. 57, no. 1, pp. 289–300, 1995
1995
-
[53]
Equivalence tests: A practical primer for t tests, correlations, and meta-analyses,
D. Lakens, “Equivalence tests: A practical primer for t tests, correlations, and meta-analyses,”Soc. Psychol. Personal. Sci., vol. 8, no. 4, pp. 355– 362, 2017
2017
-
[54]
Agresti,Analysis of Ordinal Categorical Data
A. Agresti,Analysis of Ordinal Categorical Data. Gainesville, FL, USA: Wiley, 2010, vol. 2
2010
-
[55]
A general and simple method for obtaining r2 from generalized linear mixed-effects models,
S. Nakagawa and H. Schielzeth, “A general and simple method for obtaining r2 from generalized linear mixed-effects models,”Methods Ecol. Evol., vol. 4, no. 2, pp. 133–142, 2013. VII. BIOGRAPHYSECTION Lukas K ¨oningreceived the B.Sc. degree and the M.Sc. degree in Automotive En...
2013
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.