REVIEW 2 major objections 4 minor 52 references
On Error Classification from Physiological Signals within Airborne Environment
T0 review · 2 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Live flight trials show EEG can detect pilot errors during real flight, at 87.8% accuracy.
desk verdict Hard-to-get flight data and a reasonable feasibility claim, but the missing correct-action control means the headline accuracy may reflect event detection as much as error detection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the error window: a tablet-based multi-tasking tool automatically logs a timestamp for every task error, and each timestamp is paired with a one-second segment of synchronized physiological data. Randomly sampled non-error windows form the other class, and standard classifiers (Random Forest, AdaBoost, Multi-layer Perceptron) are trained separately for EEG, eye-tracking, and ECG. The error window is what lets the paper compare airborne and laboratory conditions on the same footing, and the one-second duration is chosen to capture the 50–500 ms error-related brain potentials established in laboratory EEG work.
What would settle it
Re-train the same classifiers with error windows versus correct-response windows, such as correct classifications or warnings acknowledged within the time limit, matched for timing and motor demands; if accuracy falls to chance, the 87.83% figure measures response-related physiology rather than error detection.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that error-related physiology transfers to the airborne environment. EEG classifiers trained on one-second windows reached 87.83% accuracy during airborne trials versus 89.23% in the laboratory, and leave-one-participant-out validation still reached 86.80%, suggesting the effect is not tied to a single pilot. Eye-tracking remained moderately informative at 82.50%, while ECG fell to 51.50%, close to chance, which the authors attribute to cardiac signals being dominated by the physical demands of flight. Accuracy stayed stable across straight-and-level and 2G conditions, with only a 2.1 percentage point drop for EEG under 2G. The authors present this as the first evidence that established laboratory error-detection approaches can translate to operational aviation environments.
Load-bearing premise
The load-bearing assumption is that each tool-logged timestamp marks a true error state and that the one-second window around it isolates error-related physiology; because the non-error class is a random sample of other windows rather than a set of known correct responses, the reported accuracy could partly reflect action-related brain, gaze, or response-timing correlates rather than error-specific processing.
Editorial extensions
If this is right
- A cockpit system could flag probable errors in real time using EEG, since the signal stays informative during actual flight and under 2G load.
- The small accuracy gap between lab and flight (89.23% vs. 87.83%) gives a concrete benchmark: airborne EEG is not a fundamentally degraded version of the lab signal.
- Eye-tracking can serve as a complementary channel: pilots with lower EEG accuracy often had higher eye-tracking accuracy, arguing for multimodal monitoring.
- ECG should be deprioritized for in-flight error detection, since its near-chance performance suggests cardiac measures in this context reflect physical exertion rather than cognitive error.
- Above 86% leave-one-participant-out accuracy for EEG suggests a general error classifier could be deployed without calibrating to each pilot individually.
Reading between the lines
- Editorial inference: the paper does not compare error windows with windows of correct responses; its non-error class is a random sample of all other windows, so part of the reported accuracy may come from detecting any task-relevant action or response rather than the error state specifically.
- Editorial inference: the data only cover post-error classification, so the same recordings could be re-windowed to test whether pre-error precursors exist, which would be needed for intervention before an error rather than after it.
- Editorial inference: the sample is nine male commercial pilots on one aircraft type, so transfer to other pilot populations, fatigue states, or aircraft remains an open question that this study does not address.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a live-flight feasibility study of classifying operator errors from physiological signals (EEG, eye tracking, ECG) in nine commercial pilots. Errors are defined via timestamps logged by the IMPACT multi-tasking tool; features are extracted from 1-second windows and classified as error vs. non-error using random forests, AdaBoost, and MLP under within-subject cross-validation and leave-one-participant-out validation. The authors report EEG accuracy of 87.83% in airborne conditions (89.23% in the lab), eye-tracking accuracy of 82.50%, and ECG accuracy near chance (51.50%), and interpret these results as the first evidence that physiological error detection can transfer to operational aviation environments.
Significance. If the reported accuracy reflects error-specific cognitive processing, the study would be a valuable first step toward in-flight error monitoring, combining a realistic operational setting (including 2G manoeuvres) with careful validation protocols (per-participant, within-subject, and leave-one-participant-out evaluations). The detailed reporting of per-participant variability and the head-to-head comparison of EEG, eye tracking, and ECG under physical load are useful contributions. However, the central interpretation depends on the classification contrast in Section 3.6, which currently does not control for correct-action windows; until that is addressed, the study demonstrates discrimination of error-timestamp windows from a random sample of background time, not error detection per se.
major comments (2)
- [3.6 and 3.3] The binary classification in Section 3.6 contrasts 1-second windows at error timestamps logged by the IMPACT Tool against 'uniformly sampled timestamps of non-error-events.' This comparison does not isolate error-specific processing. The non-error class includes rest periods, low-engagement intervals, and — critically — windows around correct task actions (e.g., correct classifications, timely alert acknowledgments). A classifier can therefore attain high accuracy by detecting generic markers of task engagement or action execution, such as motor potentials from button presses, alert-evoked potentials, or gaze shifts, rather than error-related neural activity. The EEG feature set in Appendix C (frequency bands, morphology, wavelets, AR coefficients) is fully capable of carrying such event-related correlates. The near-chance ECG result does not resolve this concern, because cardiac responses are slower and strongly affected by physical load. Please re-analyze with a control condition matched for task events (e.g., correct-action windows) or, if that is not feasible, explicitly reframe the claims as task-event discrimination rather than error detection.
- [Abstract, Section 5] The abstract and conclusion claim 'the first evidence that physiological error detection can translate effectively to operational aviation environments.' Given the control-condition issue above, this claim is not supported by the reported analyses. The study provides evidence that EEG can separate error-timestamp windows from background activity in flight, which is a necessary but not sufficient step toward demonstrating error detection. The conclusions should be tempered unless a matched correct-response analysis is provided.
minor comments (4)
- [Section 3.6] The manuscript does not report the number of error events per participant or condition, nor the total number of windows used for training. Reporting these counts (and the class-balance ratio) would help readers assess the stability of the accuracy estimates and the per-participant variability in Figure 4.
- [Section 3.1 and Appendix A] The power analysis is described with inconsistent parameters: Section 3.1 states a medium effect size f = 0.35, while Appendix A states an odds ratio of 1.5 for logistic regression. Please align these descriptions and specify the primary outcome measure used for the power calculation.
- [Section 4] The statement that the 2G degradation is 'p > .001' is not a standard way to report a non-significant result; please report the test statistic, degrees of freedom, and exact p-value, and consider a correction for multiple comparisons across modalities and environments.
- [Section 3.3] The IMPACT Tool error definitions include heterogeneous events (misclassifications, missed responses, targets moving off-screen, and alert timeouts). These may have different physiological signatures; the paper would benefit from reporting whether results are stable when each error type is analyzed separately, or at least acknowledging this heterogeneity as a limitation.
Circularity Check
No circular derivation: EEG accuracy is an empirical, held-out classification result, not a fitted constant; self-citations are not load-bearing.
full rationale
The central claim is an empirical generalization from supervised classification. Error timestamps are logged by the IMPACT Tool independently of the physiological features; features (EEG bands, morphology, wavelets, AR coefficients) are computed from raw signals; classifiers are trained and tested on disjoint folds ('5-fold cross-validation while maintaining subject separation between folds' and 'leave-one-participant-out validation'). The reported 87.83% airborne EEG accuracy is therefore a prediction on unseen windows, not a parameter refit or a quantity defined by the labels. The only self-citations are for preprocessing (e.g., [19] for EEG re-referencing) and generic classifier selection ([26,29]); none of these supplies the load-bearing argument or forbids alternatives. The skeptic's point that the non-error class is a uniform sample of all non-error windows, rather than correct-response windows, is a construct-validity concern about what is being learned, not a circularity: the paper's stated operationalization is exactly binary classification of IMPACT-logged error timestamps versus other timestamps, and the accuracy numbers honestly measure separability under that operationalization. No equation or fitted input reduces the result to its own definition, so the circularity score is minimal.
Assumptions & free parameters
free parameters (2)
- classifier hyperparameters and model selection (RF/AdaBoost/MLP) =
not reported
- error/non-error window balance ratio =
not reported (described as 'uniformly selected')
assumptions (3)
- domain assumption IMPACT Tool error timestamps are valid ground-truth labels of human error.
- domain assumption Error-related physiology is captured in 1-second windows starting at (or near) the error timestamp.
- domain assumption Non-error windows sampled uniformly from non-error events form a fair comparison set.
Cite this review
Pith. "Pith review of On Error Classification from Physiological Signals within Airborne Environment." pith.science (2026). https://pith.science/paper/EOJOTG3Z
@misc{pith2026250412769,
author = {Pith},
title = {Pith review of: On Error Classification from Physiological Signals within Airborne Environment},
year = {2026},
howpublished = {\url{https://pith.science/paper/EOJOTG3Z}},
note = {Machine review of arXiv:2504.12769}
}
read the original abstract
Human error remains a critical concern in aviation safety, contributing to 70-80% of accidents despite technological advancements. While physiological measures show promise for error detection in laboratory settings, their effectiveness in dynamic flight environments remains underexplored. Through live flight trials with nine commercial pilots, we investigated whether established error-detection approaches maintain accuracy during actual flight operations. Participants completed standardized multi-tasking scenarios across conditions ranging from laboratory settings to straight-and-level flight and 2G manoeuvres while we collected synchronized physiological data. Our findings demonstrate that EEG-based classification maintains high accuracy (87.83%) during complex flight manoeuvres, comparable to laboratory performance (89.23%). Eye-tracking showed moderate performance (82.50\%), while ECG performed near chance level (51.50%). Classification accuracy remained stable across flight conditions, with minimal degradation during 2G manoeuvres. These results provide the first evidence that physiological error detection can translate effectively to operational aviation environments.
Figures
Reference graph
Works this paper leans on
-
[1]
Dorina-Marcela Ancau, Mircea Ancau, and Mihai Ancau. 2022. Deep-learning online EEG decoding brain-computer interface using error-related potentials recorded with a consumer-grade headset. Biomedical Physics & Engineering Express 8, 2 (2022), 025006
work page 2022
-
[2]
Jackson Beatty. 1982. Task-evoked pupillary responses, processing load, and the structure of processing resources. Psychological bulletin 91, 2 (1982), 276
work page 1982
-
[3]
Ricardo Chavarriaga and José del R Millán. 2010. Learning from EEG error-related potentials in noninvasive brain-computer interfaces. IEEE transactions on neural systems and rehabilitation engineering 18, 4 (2010), 381–388
work page 2010
-
[4]
Federico Chella, Vittorio Pizzella, Filippo Zappasodi, and Laura Marzetti. 2016. Calibration of a spline-based high-density EEG system: A comparison between MEG and EEG data. Frontiers in neuroinformatics 10 (2016), 13
work page 2016
-
[5]
Frédéric Dehais, Julia Behrend, Vsevolod Peysakhovich, Mickaël Causse, and Christopher D Wickens. 2017. Pilot flying and pilot monitoring’s aircraft state awareness during go-around execution in aviation: A behavioral and eye tracking study. The International Journal of Aerospace Psychology 27, 1-2 (2017), 15–28
work page 2017
-
[6]
Frédéric Dehais, Alex Lafont, Raphaëlle Roy, and Stephen Fairclough. 2020. A neuroergonomics approach to mental workload, engagement and human perfor- mance. Frontiers in neuroscience 14 (2020), 268
work page 2020
-
[7]
Andrew Duchowski and Andrew Duchowski. 2007. Eye tracking techniques. Eye tracking methodology: Theory and practice (2007), 51–59
work page 2007
-
[8]
Stephen H Fairclough. 2009. Fundamentals of physiological computing. Interact- ing with computers 21, 1-2 (2009), 133–145
work page 2009
Show all 52 references
-
[9]
Michael Falkenstein, Joachim Hohnsbein, Jörg Hoormann, and Ludger Blanke
-
[10]
William J Gehring, Brian Goss, Michael GH Coles, David E Meyer, and Emanuel Donchin. 1993. A neural system for error detection and compensation. Psycho- logical science 4, 6 (1993), 385–390
1993
-
[11]
Alan Gevins, Michael E Smith, Linda McEvoy, and Daphne Yu. 1997. High- resolution EEG mapping of cortical activation related to working memory: effects of task difficulty, type of processing, and practice. Cerebral cortex (New York, NY:
1997
-
[12]
Soo-Yeon Han, No-Sang Kwak, Taegeun Oh, and Seong-Whan Lee. 2020. Clas- sification of pilots’ mental states using a multimodal deep learning network. Biocybernetics and Biomedical Engineering 40, 1 (2020), 324–336
2020
-
[13]
7, 4 (1997), 374–385
1997
-
[14]
Hastie, Saharon Rosset, Ji Zhu, and Hui Zou
Trevor J. Hastie, Saharon Rosset, Ji Zhu, and Hui Zou. 2009. Multi-class AdaBoost *. Statistics and Its Interface 2 (2009), 349–360. https://api.semanticscholar.org/ CorpusID:11803458
2009
-
[15]
RI Harris and SE Steare. 2006. A meta-analysis of ECG data from healthy male vol- unteers: diurnal and intra-subject variability, and implications for planning ECG assessments and statistical analysis in clinical pharmacology studies. European journal of clinical pharmacology ...
2006
-
[16]
Katrina Hinde, Graham White, and Nicola Armstrong. 2021. Wearable devices suitable for monitoring twenty four hour heart rate variability in military popu- lations. Sensors 21, 4 (2021), 1061
2021
-
[17]
Christoph S Herrmann, Daniel Strüber, Randolph F Helfrich, and Andreas K Engel. 2016. EEG oscillations: from correlation to causality. International Journal of Psychophysiology 103 (2016), 12–21
2016
-
[18]
Kenneth Holmqvist, Richard Andersson, Richard Dewhurst, Halszka Jarodzka, Joost Van de Weijer, et al. 2011. Eye tracking: A comprehensive guide to methods and measures. oup Oxford
2011
-
[19]
Jean-Michel Hoc. 2000. From human–machine interaction to human–machine cooperation. Ergonomics 43, 7 (2000), 833–843
2000
-
[20]
Wolfgang Klimesch. 1999. EEG alpha and theta oscillations reflect cognitive and memory performance: a review and analysis. Brain research reviews 29, 2-3 (1999), 169–195
1999
-
[21]
Kunjira Kingphai and Yashar Moshfeghi. 2021. On EEG preprocessing role in deep learning effectiveness for mental workload classification. In Human Mental Workload: Models and Applications: 5th International Symposium, H-WORKLOAD 2021, Virtual Event, November 24–26, 2021, Proce...
2021
-
[22]
Steven J Luck. 2014. An introduction to the event-related potential technique . MIT press
2014
-
[23]
Guohua Li, Susan P Baker, Jurek G Grabowski, and George W Rebok. 2001. Factors associated with pilot error in aviation crashes. A viation, space, and environmental medicine 72, 1 (2001), 52–58
2001
-
[24]
Marek Malik. 1996. Heart rate variability: Standards of measurement, physio- logical interpretation, and clinical use: Task force of the European Society of Cardiology and the North American Society for Pacing and Electrophysiology. Annals of Noninvasive Electrocardiology 1, 2...
1996
-
[25]
Päivi Majaranta and Andreas Bulling. 2014. Eye tracking and eye-based human– computer interaction. In Advances in physiological computing . Springer, 39–65
2014
-
[26]
Niall McGuire and Yashar Moshfeghi. 2023. What song am I thinking of?. In International Conference on Machine Learning, Optimization, and Data Science . Springer, 418–432
2023
-
[27]
Daniel Martinez-Marquez, Sravan Pingali, Kriengsak Panuwatwanich, Rodney A Stewart, and Sherif Mohamed. 2021. Application of eye tracking technology in aviation, maritime, and construction industries: A systematic review. Sensors 21, 13 (2021), 4289
2021
-
[28]
Niall McGuire and Yashar Moshfeghi. 2024. Prediction of the Realisation of an Information Need: An EEG Study. In Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. https://api.semanticscholar. org/CorpusID:270391707
2024
-
[29]
Niall McGuire and Yashar Moshfeghi. 2024. DEEPER: Dense Electroencephalog- raphy Passage Retrieval. arXiv preprint arXiv:2412.06695 (2024)
2024 arXiv
-
[30]
Tiago Palma Pagano, Rafael Bessa Loureiro, Fernanda Vitoria Nascimento Lisboa, Rodrigo Matos Peixoto, Guilherme A. S. Guimarães, Gustavo Oliveira Ramos Cruz, Maira M. Araujo, Lucas Lisboa dos Santos, Marco A. S. Cruz, Ewerton L. S. Oliveira, Ingrid Winkler, and Erick Giovani S...
2023
-
[31]
Niall McGuire and Yashar Moshfeghi. 2024. Prediction of the realisation of an information need: an EEG study. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2584– 2588
2024
-
[32]
Raja Parasuraman, Thomas B Sheridan, and Christopher D Wickens. 2000. A model for types and levels of human interaction with automation. IEEE Transac- tions on systems, man, and cybernetics-Part A: Systems and Humans 30, 3 (2000), 286–297
2000
-
[33]
Mahesh Pal. 2005. Random forest classifier for remote sensing classifica- tion. International Journal of Remote Sensing 26 (2005), 217 – 222. https: //api.semanticscholar.org/CorpusID:131528125
2005
-
[34]
Hongquan Qu, Xueying Gao, and Liping Pang. 2021. Classification of mental workload based on multiple features of ECG signals. Informatics in Medicine Unlocked 24 (2021), 100575
2021
-
[35]
Mastorakis
Mariusconstantin Popescu, Valentina Emilia Balas, Liliana Perescu-Popescu, and Nikos E. Mastorakis. 2009. Multilayer perceptron and neural networks. WSEAS Transactions on Circuits and Systems archive 8 (2009), 579–588. https://api. semanticscholar.org/CorpusID:13978378
2009
-
[36]
Jennifer M Riley, Laura D Strater, Sheryl L Chappell, Erik S Connors, and Mica R Endsley. 2016. Situation awareness in human-robot interaction: Challenges and user interface requirements. In Human-Robot Interactions in Future Military Operations. CRC Press, 171–192
2016
-
[37]
James Reason. 1990. Human error. Cambridge university press
1990
-
[38]
Thomas B Sheridan. 2016. Human–robot interaction: status and challenges. Human factors 58, 4 (2016), 525–532
2016
-
[39]
G. Sabine. 2022. Interactive Measures of Performance and Assessment of Cogni- tive Tasks (IMPACT) Tool. Technical Report DSTL/TR139745. Defence Science Technology Laboratory. On Error Classification from Physiological Signals within Airborne Environment CHI EA ’25, April 26-Ma...
2022
-
[40]
Michal Teplan. 2002. Fundamentals of EEG measurement. Measurement science review 2, 2 (2002), 1–11
2002
-
[41]
A Stuiver, D De Waard, KA Brookhuis, C Dijksterhuis, B Lewis-Evans, and LJM Mulder. 2012. Short-term cardiovascular responses to changing task demands. International Journal of Psychophysiology 85, 2 (2012), 153–160
2012
-
[42]
Chi Thanh Vi, Izdihar Jamil, David Coyle, and Sriram Subramanian. 2014. Error related negativity in observing interactive tasks. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Toronto, Ontario, Canada) (CHI ’14). Association for Computing Machin...
2014
-
[43]
Chi Vi and Sriram Subramanian. 2012. Detecting error-related negativity for interaction design. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Austin, Texas, USA)(CHI ’12). Association for Computing Ma- chinery, New York, NY, USA, 493–502. https...
2012
-
[44]
Douglas A Wiegmann and Scott A Shappell. 2017. A human error approach to aviation accident analysis: The human factors analysis and classification system . Routledge
2017
-
[45]
Christopher D Wickens, Alan Stokes, Barbara Barnett, and Fred Hyman. 2017. The effects of stress on pilot judgment in a MIDIS simulator. In Decision Making in A viation. Routledge, 387–408
2017
-
[46]
Glenn F Wilson, George A Reis, and Lloyd D Tripp. 2005. EEG correlates of G-induced loss of consciousness. A viation, space, and environmental medicine 76, 1 (2005), 19–27
2005
-
[47]
Glenn F Wilson and John A Caldwell Jr. 2002. Cardiac and eye activity correlates of sleep loss in helicopter pilots. InProceedings of the Human Factors and Ergonomics Society Annual Meeting, Vol. 46. SAGE Publications Sage CA: Los Angeles, CA, 126–129
2002
-
[48]
Dezhong Yao. 2001. A theoretical study of the effects of volume conductor on EEG and ECoG. IEEE Transactions on Biomedical Engineering 48, 1 (2001), 87–96
2001
-
[49]
C Wirth, PM Dockree, S Harty, E Lacey, and M Arvaneh. 2019. Towards error categorisation in BCI: single-trial EEG classification between different errors. Journal of neural engineering 17, 1 (2019), 016008
2019
-
[50]
Thorsten O Zander and Christian Kothe. 2011. Towards passive brain–computer interfaces: applying brain–computer interface technology to human–machine systems in general. Journal of neural engineering 8, 2 (2011), 025005. A POWER ANALYSIS To determine the required sample size f...
2011
-
[51]
Yei-Yu Yeh and Christopher D Wickens. 1988. Dissociation of performance and subjective measures of workload. Human factors 30, 1 (1988), 111–120
1988
-
[1991]
Effects of crossmodal divided attention on late ERP components. II. Er- ror processing in choice reaction tasks. Electroencephalography and clinical neurophysiology 78, 6 (1991), 447–455
1991
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.