REVIEW 4 major objections 5 minor 32 references
Exploring Eye Tracking to Detect Cognitive Load in Complex Virtual Reality Training
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper reports that mean pupil dilation and fixation duration, fed into a five-layer neural network, can separate high from low self-reported mental workload in a 26-minute VR cold-spray assembly task with 84% accuracy.
desk verdict A real pilot study with an honest limitations section, but the 0.84 accuracy is not interpretable as written because the sliding-window evaluation split is underspecified and may leak participant identity. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a supervised classifier operating on a sliding window of eye-tracking signals. Raw gaze data are first denoised with a Fast Fourier Transform and normalized; a commercial eye-tracking processing pipeline, using the Velocity-Threshold Identification (I-VT) filter with a 60 ms minimum, converts the signal into fixation events. For each fixation, the feature vector is the fixation duration and the mean pupil dilation across both eyes. A sliding window of 2000 samples chunks the continuous stream into instances, and the binary target is the participant's NASA-TLX mental-demand group (low=1-4, high=5-7), assigned to every window from that participant. The architecture—an MLP with five hidden layers and hyperbolic tangent activations, trained with Adam at learning rate 0.00001 for 500 epochs with batch size 256—maps these windows to a high/low workload class; a Random Forest tuned with GridSearchCV serves as the comparison model.
What would settle it
Re-run the same pipeline with a strict participant-independent split (e.g., leave-one-participant-out: train on 18 participants, test on the remaining one, repeat) and check whether the mean accuracy stays near 0.84; if it falls toward the majority-class baseline, the reported accuracy was driven by participant-specific leakage rather than a general cognitive-load signal.
Extended reading notes
Core claim
The central claim is that eye-tracking features carry enough information about mental workload to separate high from low cognitive load in a complex spatiotemporal VR task, even with a modest number of users. In this study, each participant performed a 12-step disassembly and 11-step assembly of a powder feeder in a cold-spray VR environment, taking about 26 minutes. The authors extracted fixation duration and the mean of both eyes' pupil dilation within each fixation using an I-VT filter, denoised and normalized the signals, and labeled each participant by the mental-demand subscale of NASA-TLX (scores 1-4 low, 5-7 high). A Multi-Layer Perceptron with five tanh hidden layers, trained with Adam at learning rate 0.00001 over 500 epochs on 2000-sample windows, achieved 0.84 accuracy and precision on the test set, with recall 0.94; a tuned Random Forest reached 0.72 accuracy. The authors interpret the MLP's performance as evidence that reliable, non-intrusive cognitive-load detection in dynamic VR training is feasible.
Load-bearing premise
The result stands or falls on whether the 0.84 test accuracy was computed on windows from participants who were not seen during training; the paper does not describe the train/test split, and every window from one participant shares the same NASA-TLX label, so overlapping participants could let the model exploit individual pupil patterns instead of cognitive load.
Editorial extensions
If this is right
- A real-time classifier using only pupil dilation and fixation duration could drive adaptive VR training that changes task difficulty or provides help when cognitive load is high.
- The MLP's 0.84 accuracy over the Random Forest's 0.72 supports non-linear models for eye-tracking-based workload prediction in dynamic tasks.
- Because the split of NASA-TLX mental-demand into low (1-4) and high (5-7) was balanced, the binary formulation ties the classifier to a practical go/no-go signal for adaptation.
- Privacy and data security must be addressed before deployment because gaze data can reveal personal traits and health information.
- Expanding features to saccade velocity, saccade amplitude, and blink rate could improve robustness, as the authors note.
Reading between the lines
- The reported 0.84 accuracy may overstate generalizable performance: because the paper does not specify whether the sliding-window test instances come from participants held out of training, the classifier could have learned participant-specific pupil patterns rather than a general workload signature; a leave-one-participant-out evaluation would settle this.
- The NASA-TLX mental-demand label is one score per participant, yet the classifier receives many windows from the same participant; modeling the label as constant per participant means the effective sample size for generalization is closer to 19 than to the number of windows, and the confidence intervals around 0.84 would be wide.
- If the participant-level generalization holds, the same pipeline could be transferred to other procedural VR training tasks (surgical, maintenance, logistics) that have similar step-by-step assembly structures.
- Combining pupil dilation with other gaze events (saccades, blinks) or physiological signals (heart rate, electrodermal activity) could make the workload estimate robust to lighting changes and individual pupil-size baselines.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports an ongoing study of eye-tracking-based cognitive load detection in a complex spatiotemporal VR task. Twenty-two participants completed a cold-spray powder-feeder assembly/disassembly simulation in VR; 19 valid eye-tracking datasets and NASA-TLX scores were retained. From each participant's data the authors extracted mean pupil dilation and fixation duration, segmented the signals into 2000-sample sliding windows, and trained MLP and random forest classifiers to predict a binary target created by splitting the NASA-TLX mental-demand scores into low (1-4) and high (5-7). The central reported result is an MLP test accuracy and precision of 0.84 (RF accuracy 0.72), which the authors interpret as demonstrating the feasibility of detecting cognitive load in VR training and motivating future adaptive training systems.
Significance. If the reported evaluation were methodologically airtight, the paper would be a useful empirical data point for adaptive VR training in advanced manufacturing, since it uses a realistic 26-minute assembly task and commercially available VR eye tracking. The authors are appropriately modest in calling the study preliminary and in discussing privacy implications. However, the central claim is not currently supported by the reported methodology: the evaluation protocol permits participant-level label leakage, and the target threshold is chosen from the same sample used for evaluation. Because the contribution is purely empirical, these issues are load-bearing. The paper does not ship code, machine-checked derivations, or preregistered analyses, so the correctness of the evaluation is the only basis for the headline accuracy figure.
major comments (4)
- [Section 2.2 and Section 3] The manuscript never states how the 2000-sample sliding windows were partitioned into training and test sets. Every window from a given participant carries the same binarized NASA-TLX score, so if windows from one participant appear in both training and test, the classifier can learn participant-specific pupil or fixation patterns rather than a generalizable cognitive-load signature. The reported MLP accuracy of 0.84 in Section 3 is therefore not interpretable unless the split is participant-independent. Please report a leave-one-participant-out or grouped cross-validation procedure, the number of windows per participant, and the class distribution.
- [Section 2.2] The low/high labels are defined by a cutoff (1-4 vs 5-7) that the authors state was chosen after analyzing the distribution of the NASA-TLX mental-demand scores from the same 19 participants. This makes the target itself dependent on the evaluation sample; selecting a cutoff that produces an even split can inflate apparent separability and weakens the claim that the classifier separates objectively defined low and high cognitive load. The threshold should be justified independently, for example from prior literature or a pre-registered criterion, and the resulting class sizes should be reported.
- [Section 3 and Table 1] The reported metrics are point estimates with no baseline comparison, confidence intervals, or class counts. With only 19 participants and a binary label, a majority-class classifier or a model exploiting participant identity can achieve non-trivial accuracy. Table 1 alone cannot rule out these alternatives. Please provide the confusion matrix, per-class precision and recall, a majority-class or chance baseline, and the variance of metrics across cross-validation folds.
- [Section 2.2 and Section 4] A single post-session NASA-TLX score is used as the label for every window in a participant's session, but the stated goal is real-time adaptive training. The model is therefore trained to predict a session-level construct from short temporal windows, and no time-resolved ground truth is available to validate the real-time claim. The authors should either temper the real-time framing or evaluate on a workload measure that varies within the session.
minor comments (5)
- [Section 3] "Tabel 1" should be "Table 1".
- [Section 2] The phrase "thecold spray" is missing a space and should read "the cold spray".
- [Section 2.2] Please specify the temporal duration of the 2000-sample sliding window given the recording rate stated in Section 2.1, and state the stride or overlap between consecutive windows.
- [References] Reference [1] contains a typographical error in the author initials; it should be corrected to "Y. V. Wong" or the intended spelling.
- [Section 2.2] The sentence "The MLP's performance was carefully monitored, and the training process was optimized to maintain high precision and recall on both the training and test sets" is vague; please describe the monitoring procedure and the overfitting checks quantitatively.
Circularity Check
The 0.84 test accuracy is not an independent prediction: the training process was optimized against the test set, and the high/low target was defined by a post-hoc split of the same sample.
-
self definitional
[Section 2.2, target-variable definition (paragraph beginning 'The target variable was the mental demand subsection...')]
"We analyzed the distributions of the NASA-TLX mental workload subsection and found that the population could be evenly split into low and high groups, with scores of 1 to 4 classified as low and 5 to 7 as high. Accordingly, we labeled participants into either a high or low group and assigned each the corresponding target variable."
The binary outcome the models are said to predict is not an independently defined cognitive-load construct; the 1-to-4 versus 5-to-7 cutoff was chosen after inspecting the distribution of the 19 participants whose eye-tracking windows are later classified. The target is therefore a sample-relative post-hoc partition of the evaluation data, and the reported accuracy is accuracy at recovering a label constructed from that same sample. This does not force the eye-feature mapping by itself, but it makes the prediction target self-referential rather than externally anchored.
-
fitted input called prediction
[Section 2.2, MLP training paragraph and Section 3, first results paragraph]
"The MLP’s performance was carefully monitored, and the training process was optimized to maintain high precision and recall on both the training and test sets. ... The MLP model, evaluated on the test dataset, achieved an accuracy and precision of 0.84, indicating strong prediction capability."
The paper explicitly says the training process was optimized to maintain high precision and recall on the test set, meaning test-set performance was an input to model selection or early stopping. The same test-set accuracy is then reported as evidence of 'strong prediction capability.' With test data used in the optimization loop, the 0.84 figure is a fitted evaluation value, not an out-of-sample prediction; the central claim therefore reduces to a score that was partially optimized rather than independently held out.
full rationale
The central numerical claim of the paper is the MLP test accuracy of 0.84 in Section 3. Two text-level features make this claim partially circular. First, the NASA-TLX mental-demand scores are binarized using a cutoff chosen after examining the distribution of the same 19 participants, so the target labels are constructed from the evaluation sample. Second, and more directly, the Methods section states that the training process was optimized to maintain high precision and recall on both the training and test sets; evaluating on a set that was used during optimization is not an independent prediction. The paper does not state how the 2000-sample sliding windows were split into training and test sets, so an additional participant-level leakage risk exists, but this is an under-specification/validity threat rather than a demonstrated circular reduction. There is no load-bearing self-citation chain: reference [18] is a companion paper about the VR environment, not about the cognitive-load result, and no uniqueness theorem is imported from the authors' prior work. The RF model's GridSearchCV with 3-fold cross-validation is a normal tuning procedure and is not circular. Because the reported test accuracy is compromised by test-set optimization and the target definition is post-hoc, the prediction claim partially reduces to its own inputs, warranting a 6 rather than a lower score. The paper remains a preliminary feasibility study and explicitly acknowledges the small sample, which limits severity but does not remove the circular evaluation issue.
Assumptions & free parameters
free parameters (4)
- NASA-TLX mental-demand low/high split threshold =
scores 1-4 = low, 5-7 = high
- sliding window size =
2000 samples
- MLP architecture and training settings =
five hidden layers, tanh activation, Adam lr=1e-5, 500 epochs, batch size 256
- RF hyperparameters =
selected via GridSearchCV with 3-fold CV
assumptions (3)
- domain assumption Pupil dilation and fixation duration are valid indicators of cognitive load
- domain assumption A single post-session NASA-TLX mental-demand score is a valid label for every 2000-sample window of the participant's 26-minute session
- domain assumption The iMotions I-VT fixation filter with 60 ms minimum duration produces valid fixation and pupil features from the Varjo VR-3
Cite this review
Pith. "Pith review of Exploring Eye Tracking to Detect Cognitive Load in Complex Virtual Reality Training." pith.science (2026). https://pith.science/paper/NOXHRIFK
@misc{pith2026241112771,
author = {Pith},
title = {Pith review of: Exploring Eye Tracking to Detect Cognitive Load in Complex Virtual Reality Training},
year = {2026},
howpublished = {\url{https://pith.science/paper/NOXHRIFK}},
note = {Machine review of arXiv:2411.12771}
}
read the original abstract
Virtual Reality (VR) has been a beneficial training tool in fields such as advanced manufacturing. However, users may experience a high cognitive load due to various factors, such as the use of VR hardware or tasks within the VR environment. Studies have shown that eye-tracking has the potential to detect cognitive load, but in the context of VR and complex spatiotemporal tasks (e.g., assembly and disassembly), it remains relatively unexplored. Here, we present an ongoing study to detect users' cognitive load using an eye-tracking-based machine learning approach. We developed a VR training system for cold spray and tested it with 22 participants, obtaining 19 valid eye-tracking datasets and NASA-TLX scores. We applied Multi-Layer Perceptron (MLP) and Random Forest (RF) models to compare the accuracy of predicting cognitive load (i.e., NASA-TLX) using pupil dilation and fixation duration. Our preliminary analysis demonstrates the feasibility of using eye tracking to detect cognitive load in complex spatiotemporal VR experiences and motivates further exploration.
Figures
Reference graph
Works this paper leans on
-
[1]
U. A. Abdurrahman, S.-C. Yeh, Y . Wong, and L. Wei. Effects of neuro- cognitive load on learning transfer using a virtual reality-based driving system. Big Data and Cognitive Computing, 5(4):54, 2021. 3
work page 2021
-
[2]
M. Ahmadi, S. W. Michalka, S. Lenzoni, M. Ahmadi Najafabadi, H. Bai, A. Sumich, B. Wuensche, and M. Billinghurst. Cognitive load measurement with physiological sensors in virtual reality during phys- ical activity. In Proceedings of the 29th ACM Symposium on Virtual Reality Software and Technology, pp. 1–11, 2023. 1, 2
work page 2023
-
[3]
A. Armougum, E. Orriols, A. Gaston-Bellegarde, C. Joie-La Marle, and P. Piolino. Virtual reality: A new method to investigate cogni- tive load during navigation. Journal of Environmental Psychology , 65:101338, 2019. 1
work page 2019
-
[4]
G. Aur ´elien. Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow. o’reilly, 2019. 3
work page 2019
- [5]
-
[6]
M. Bodaghi, M. Hosseini, and R. Gottumukkala. A multimodal in- termediate fusion network with manifold learning for stress detection. arXiv preprint arXiv:2403.08077, 2024. 1
arXiv 2024
-
[7]
S. Choi, K. Jung, and S. D. Noh. Virtual reality applications in manu- facturing industries: Past research, present findings, and future direc- tions. Concurrent Engineering, 23(1):40–63, 2015. 1
work page 2015
- [8]
Show all 32 references
-
[9]
H. Gao, L. Hasenbein, E. Bozkir, R. G ¨ollner, and E. Kasneci. Ex- ploring gender differences in computational thinking learning in a vr classroom: Developing machine learning models using eye-tracking data and explaining the models. International Journal of Artificial Intelli...
2023
-
[10]
Gao and E
H. Gao and E. Kasneci. Exploring eye tracking as a measure for cogni- tive load detection in vr locomotion. In Proceedings of the 2024 Sym- posium on Eye Tracking Research and Applications, pp. 1–3, 2024. 1, 2
2024
-
[11]
S. G. Hart. Nasa-task load index (nasa-tlx); 20 years later. In Pro- ceedings of the human factors and ergonomics society annual meet- ing, vol. 50, pp. 904–908. Sage publications Sage CA: Los Angeles, CA, 2006. 2
2006
-
[12]
imotions lab
iMotions. imotions lab. https://imotions.com/products/ imotions-lab/, 2024. Last accessed: 07/26/2024. 2
2024
-
[13]
J. L. Kr ¨oger, O. H.-M. Lutz, and F. M ¨uller. What does your gaze reveal about you? on the privacy implications of eye tracking. InIFIP International Summer School on Privacy and Identity Management , pp. 226–241. Springer, 2020. 3
2020
-
[14]
J. Y . Lee, N. de Jong, J. Donkers, H. Jarodzka, and J. J. van Merri¨enboer. Measuring cognitive load in virtual reality training via pupillometry. IEEE Transactions on Learning Technologies, 2023. 2
2023
-
[15]
Liu, K.-A
J.-C. Liu, K.-A. Li, S.-L. Yeh, and S.-Y . Chien. Assessing perceptual load and cognitive load by fixation-related information of eye move- ments. Sensors, 22(3):1187, 2022. 2
2022
-
[16]
R. N. M eghanathan, C. van Leeuwen, and A. R. Nikolaev. Fixation duration surpasses pupil size as a measure of memory load in free viewing. Frontiers in human neuroscience, 8:1063, 2015. 2
2015
-
[17]
Miles, M
G. Miles, M. Smith, N. Zook, and W. Zhang. Em-cogload: An inves- tigation into age and cognitive load detection using eye tracking and deep learning. Computational and Structural Biotechnology Journal, 24:264–280, 2024. 2
2024
-
[18]
Nasri, U
M. Nasri, U. Narayan, M. Feyyaz Sonbudak, A. Simonson, M. Chiu, J. Donati, M. Sivak, M. Kosa, and C. Harteveld. Designing a virtual reality training apprenticeship for cold spray advanced manufacturing. arXiv e-prints, pp. arXiv–2411, 2024. 2
2024
-
[19]
F. Paas, A. Renkl, and J. Sweller. Cognitive load theory and in- structional design: Recent developments. Educational psychologist, 38(1):1–4, 2003. 1
2003
-
[20]
Papyrin, V
A. Papyrin, V . Kosarev, S. Klinkov, A. Alkhimov, and V . M. Fomin. Cold spray technology. Elsevier, 2006. 2
2006
-
[21]
D. D. Salvucci and J. H. Goldberg. Identifying fixations and saccades in eye-tracking protocols. In Proceedings of the 2000 symposium on Eye tracking research & applications, pp. 71–78, 2000. 2
2000
-
[22]
Shojaeizadeh, S
M. Shojaeizadeh, S. Djamasbi, R. C. Paffenroth, and A. C. Trapp. Detecting task demand via an eye tracking machine learning system. Decision Support Systems, 116:91–101, 2019. 2
2019
-
[23]
Skaramagkas, E
V . Skaramagkas, E. Ktistakis, D. Manousos, N. S. Tachos, E. Kazantzaki, E. E. Tripoliti, D. I. Fotiadis, and M. Tsiknakis. Cog- nitive workload level estimation based on eye tracking: A machine learning approach. In 2021 IEEE 21st International Conference on Bioinformatics an...
2021
-
[24]
A. D. Souchet, S. Philippe, D. Lourdeaux, and L. Leroy. Measuring visual fatigue and cognitive load via eye tracking while learning with virtual reality head-mounted displays: A review. International Jour- nal of Human–Computer Interaction, 38(9):801–824, 2022. 1, 2
2022
-
[25]
Szczepaniak, M
D. Szczepaniak, M. Harvey, and F. Deligianni. Predictive modelling of cognitive workload in vr: An eye-tracking approach. In Proceedings of the 2024 Symposium on Eye Tracking Research and Applications , pp. 1–3, 2024. 1, 2
2024
-
[26]
Technologies
U. Technologies. Unity. https://unity.com/, 2024. 2
2024
-
[27]
Thomay, A
C. Thomay, A. Fermitsch, J. Fessler, P. Garatva, B. Gollan, A. Katha- rina Lietz, M. Matscheko, and M. Wagner. Towards cognitive load- based decision making in vr training. In 2023 IEEE 2nd Interna- tional Conference on Cognitive Aspects of Virtual Reality (CVR) , pp. 000023–0...
2023
-
[28]
Thomay, A
C. Thomay, A. Fermitsch, J. Fessler, P. Garatva, B. Gollan, A. K. Li- etz, M. Matscheko, and M. Wagner. Towards cognitive load-based decision making in vr training. In 2023 IEEE 2nd International Con- ference on Cognitive Aspects of Virtual Reality (CVR) , pp. 000023– 000028. ...
2023
-
[29]
Vulpe-Grigorasi
A. Vulpe-Grigorasi. Multimodal machine learning for cognitive load based on eye tracking and biosensors. InProceedings of the 2023 Sym- posium on Eye Tracking Research and Applications, pp. 1–3, 2023. 2
2023
-
[30]
Y . Yin, C. Juan, J. Chakraborty, and M. P. McGuire. Classification of eye tracking data using a convolutional neural network. In 2018 17th IEEE International Conference on Machine Learning and Appli- cations (ICMLA), pp. 530–535. IEEE, 2018. 3
2018
-
[31]
T. O. Zander and C. Kothe. Towards passive brain–computer in- terfaces: applying brain–computer interface technology to human– machine systems in general. Journal of neural engineering , 8(2):025005, 2011. 1
2011
-
[32]
Zhang, J
L. Zhang, J. Wade, D. Bian, J. Fan, A. Swanson, A. Weitlauf, Z. War- ren, and N. Sarkar. Cognitive load measurement in a virtual reality- based driving system for autism intervention. IEEE transactions on affective computing, 8(2):176–189, 2017. 3
2017
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.