{"id":"5863ce44-c222-473d-95a9-52d4c6d7eeef","arxiv_id":"2412.16847","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review finds that wearable sensors combined with AI and multimodal fusion can detect fatigue, but the review's own numbers and meta-analysis claim are unreliable.","lead":"This paper reviews 150 studies on using wearable sensors and AI to detect fatigue, following a PRISMA-style search. It concludes that combining physiological signals like ECG and EEG with machine learning can support real-time fatigue monitoring in workplaces and other safety-critical settings.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on studies whose ground-truth labels are not validated fatigue measures; Tables 1 and 2 treat drowsiness, stress, and emotion detection as fatigue detection.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the validity of ground-truth fatigue labels. My independent reading of the manuscript confirms that this is the single most consequential issue for the central claim. The conclusion that AI plus wearables 'dramatically improved the precision... of fatigue detection systems' depends entirely on the accuracy figures in Tables 1 and 2. If those figures measure drowsiness, stress, or emotion rather than fatigue, the conclusion is not merely overstated; it is not about fatigue at all. The manuscript itself provides the ammunition for this critique in Sections 1.2 and 7.5, and the included studies and datasets repeatedly use proxy constructs such as yawning, closed eyes, sleepiness, and calm/distress classification. The PRISMA count inconsistencies and unsupported meta-analysis label are real methodological flaws, but they are secondary: even a perfectly executed systematic review would fail to support the central claim if the underlying labels are not fatigue. Therefore, I agree with the reader's REJECT verdict and recommend no change to that verdict. The concrete test proposed would settle the concern by showing, in a reproducible way, which label classes drive the high accuracy numbers.","tokens_in":46342,"tokens_out":3001,"duration_ms":27411,"concrete_test":"Build a label-audit spreadsheet from Tables 1 and 2: for each included primary study, fetch the paper and classify the ground-truth label as (1) validated fatigue questionnaire, (2) unvalidated self-report fatigue, (3) proxy state (drowsiness, sleepiness, stress, emotion, PVT/performance), or (4) expert/behavioral rating. Then recompute the distribution of reported accuracies by label class. If the median accuracy for class (3) is above 90% while class (1) studies are few, small-sample, or lower-accuracy, the review's central claim collapses to 'AI detects proxy states,' and the conclusion should be rewritten or the review rejected as insufficiently supporting fatigue detection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and conclusion claim that wearables plus AI provide accurate real-time fatigue monitoring. The load-bearing evidence is the accuracy numbers in Tables 1 and 2, but those numbers are only interpretable if the predicted labels actually are fatigue. Section 1.2 concedes there is no standard definition of fatigue and that physiological measures are confounded by stress, physical activity, and sleep quality; Section 7.5 concedes that patient-reported fatigue is subjective and confounded. Many Table 1 entries do not predict fatigue at all: references [100], [101], [103], [139], [45], and [72] target drowsiness or sleepiness; [114] classifies calm versus distress; [74] combines stress, fatigue, and drowsiness into one multi-class problem; [127] uses PVT reaction time, an alertness/performance proxy, as ground truth; and the YawDD, CEW, and DROZY datasets are yawning, eye-closure, and sleep-induction datasets. If the models are learning eye closure, blink rate, stress-related HRV changes, or sleepiness, the reported 90-99% accuracies do not support the conclusion that fatigue is detected. The review never audits or re-labels these ground truths, and no meta-analytic adjustment for label construct is attempted despite Section 2.2.3 claiming a 'structured meta-analysis.' Without a label audit, the central claim is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is presented as a PRISMA-guided systematic review of wearable and AI-based fatigue monitoring. It searches four databases, reports screening 393 records, and claims to include 150 studies in a 'structured meta-analysis.' It surveys physiological modalities (ECG, EMG, EEG, PPG, EDA, IMU, EOG, hybrid), catalogs datasets, tabulates reported accuracies from individual studies, discusses research challenges, and concludes that AI-powered wearables with multi-source data fusion provide accurate real-time fatigue monitoring for safety-critical settings.","tokens_in":46554,"tokens_out":4304,"duration_ms":35917,"significance":"If the central claim were sound, the review would be a useful map of an applied area with safety implications: it assembles a broad corpus, covers wearable form factors and signal modalities, and identifies real deployment gaps such as real-time data access, ergonomics, explainability, and edge computing. The paper's own discussion of unresolved definitional and validation problems is candid. However, the evidence base as presented cannot support the advertised conclusion: the review does not actually perform meta-analysis, the screening counts do not reconcile, and, most importantly, the included studies' outcome labels are not audited against the fatigue construct. The contribution is therefore currently a descriptive catalogue rather than a validated synthesis.","major_comments":[{"comment":"The PRISMA flowchart contains arithmetic inconsistencies that prevent reproduction of the study selection. The flowchart reports 393 records screened, 324 abstracts reviewed, and 24 abstracts excluded, yet the next stage lists 252 full-text articles; 393 minus 24 equals 369, and 324 minus 24 equals 300, not 252. Similarly, 252 full-text articles minus 72 excluded equals 180, not the 150 studies reported as included. These numbers must be reconciled or the selection procedure cannot be verified.","section":"Section 2.2.2, Figure 5"},{"comment":"The manuscript labels its synthesis a 'structured meta-analysis' and shows 'Studies included in quantitative synthesis (meta-analysis)' in the flowchart, but no meta-analytic methods or results are presented. There are no pooled effect sizes, no heterogeneity statistics, no risk-of-bias assessment, and no meta-analytic model. The results are narrative summaries plus per-study performance tables. The label should be corrected to 'systematic review without meta-analysis,' or the missing quantitative synthesis must be added.","section":"Section 2.2.3, Figure 5"},{"comment":"The outcome construct is not validated. Several tabulated studies predict drowsiness, sleepiness, stress, or calm/distress rather than fatigue: references [100], [101], [103], [139], [45], and [72] target drowsiness or sleepiness; [114] classifies calm versus distress; [74] merges stress, fatigue, and drowsiness into one multi-class problem; [127] uses Psychomotor Vigilance Task reaction time as ground truth; and the YawDD, CEW, and DROZY datasets are yawning, eye-closure, and sleep-induction datasets. Because the review never re-labels or audits these outcomes against the fatigue construct, the reported 60-99% accuracies do not establish that fatigue per se is detected. This directly undermines the abstract's and conclusion's claim about the precision of fatigue detection systems.","section":"Tables 1 and 2, with Sections 1.2 and 3"},{"comment":"The manuscript itself concedes that patient-reported fatigue is subjective and confounded and that no standard definition of fatigue exists, but this concession is not carried into the synthesis. No sensitivity analysis, label-quality stratification, or construct-level meta-regression is provided despite Section 2.2.3 implying such an analysis. Without this, the review cannot separate true fatigue detection from detection of correlated states such as drowsiness, stress, or physical activity change, which is the central interpretive risk of the paper.","section":"Sections 7.5 and 1.2"},{"comment":"The performance evidence is largely based on very small samples and incomplete reporting. Many rows report accuracies from 5 to 64 participants, and several entries have N/A for sample size or feature count (e.g., rows for [71], [103], [110], [127], [128], and [141]). No confidence intervals, cross-validation details, or external validation results are systematically provided. This is not fatal by itself, but combined with the label-construct issue it makes the quantitative claims in the abstract and conclusion disproportionately strong relative to the evidence.","section":"Table 1"}],"minor_comments":[{"comment":"The phrase 'a person level of exhaustion' should read 'a person's level of exhaustion.'","section":"Abstract"},{"comment":"The caption says 'Common techniques and frameworks used in [68] for fatigue measurement and monitoring,' but the figure displays a Borg scale; the caption should describe the Borg CR10 scale instead.","section":"Figure 7"},{"comment":"The sentence 'Authors in [75] [88] [84] [82] [81] [86] [85] [43] [87] [35] [76] [83] [89] have used EEG signals and AI methods' appears in the ECG-based methods subsection, and the cited studies are ECG-based; 'EEG' should be corrected to 'ECG' or the sentence should be reworked to match the subsection.","section":"Section 6.1"},{"comment":"The opening sentence is duplicated verbatim: 'When it comes to wearable devices, ergonomics and comfort are of utmost significance, particularly if they are to be worn for' appears twice in consecutive lines.","section":"Section 7.2"},{"comment":"The paragraph beginning 'The study's constraints were noted...' and the subsequent paragraph about wrist-worn sensors contain a long, unmarked methodological passage that is not clearly connected to the stated topic of edge computing for wearables; this passage should be removed or explicitly framed as a case study.","section":"Section 8.5"},{"comment":"The notation 'N /A' should be normalized to 'N/A' throughout, and 'AUD = 0.9' in the row for reference [131] should be labeled 'AUC.'","section":"Table 1"},{"comment":"The text refers to 'extended short-term memory networks (LSTM)'; the correct term is 'long short-term memory networks.'","section":"Section 6.2"}],"recommendation":"reject","confidential_remarks":"In my view the construct-validity problem is structural: the corpus conflates fatigue with drowsiness, sleepiness, and stress, so the review cannot be repaired by local edits. If the authors were to reframe the contribution as a scoping review of related alertness states and correct the PRISMA accounting, that would be a different paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the reject verdict is fair. The paper does one useful thing—it pulls together a lot of studies on wearables and AI for fatigue across ECG, EEG, EMG, PPG, EDA, IMU, and hybrid setups—but it is not a reliable synthesis. The PRISMA flowchart is internally inconsistent: 393 records screened, 324 abstracts, 24 excluded leaves 300, not 252 full-text papers; and 252 minus 72 excluded leaves 180, not 150. Those numbers are checkable, and they don't reconcile. Calling the result a 'structured meta-analysis' makes it worse: there are no pooled estimates, no heterogeneity stats, no risk-of-bias assessment. It's a tabulated narrative review.\n\nThe deeper problem is construct validity. The paper admits in Section 1.2 there is no standard definition of fatigue, and Section 7.5 concedes patient-reported fatigue is subjective and confounded. Yet Tables 1 and 2 treat drowsiness, sleepiness, and even calm-vs-distress as fatigue. Reference [45] is explicitly about driver drowsiness; [100] and [103] are drowsiness detection; [114] classifies calm and distress; [74] lumps stress, fatigue, and drowsiness into one multi-class problem. Ground truth in [127] is PVT reaction time, an alertness proxy, not fatigue. If the models are learning eye closure, blink rate, or stress-related HRV, the 90-99% accuracies don't support the paper's conclusion that wearable AI detects fatigue.\n\nCredit where it's due: the scope is broad, the tables are extensive (even if not trustworthy for the reasons above), and the challenges/future directions sections (real-time data, ergonomics, XAI, uncertainty, edge computing) are reasonable and reasonably complete. The authors also do flag the definitional problem themselves and list confounders; they just don't follow through.\n\nWho is this for? A reader who wants a quick map of which modalities have been tried and which ML methods are common could skim Section 6 and the tables. Nobody should cite the accuracy numbers as evidence for fatigue detection without checking the primary studies.\n\nMy recommendation: send it to peer review, but with a clear expectation of major revision. The topic matters, the flaws are hard to ignore but are fixable in principle: correct the PRISMA counts, retract the meta-analysis claim, and audit every included study's label—drop or reclassify the drowsiness/stress papers. If the authors can't do that, reject. If they can, this could become a serviceable, if not groundbreaking, review.","headline":"A broad but sloppy review that over-claims: the included studies often detect drowsiness or stress, not fatigue, and the PRISMA counts don't add up.","tokens_in":47117,"tokens_out":3361,"would_cite":false,"duration_ms":27023,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI wearables can track fatigue in real time","keywords":["fatigue monitoring","wearable sensors","artificial intelligence","machine learning","deep learning","physiological signals","multimodal fusion","real-time health monitoring"],"falsifier":"A field study equipping several hundred shift workers with multisignal wearables and probing fatigue with an objective benchmark, such as psychomotor vigilance task reaction times at random times, could settle whether the reported lab accuracies generalize. If the AI models' fatigue predictions match the probes better than chance, the central claim holds; if not, the review's conclusion would need to be weakened.","tokens_in":46130,"feed_emoji":"⌚","tokens_out":5545,"duration_ms":44987,"temperature":0.7,"pith_summary":"This review asks whether consumer-grade wearables combined with machine and deep learning can deliver accurate, real-time fatigue monitoring. The authors argue that fusing multiple physiological signals—such as ECG, EEG, EMG, and PPG—and analyzing them with AI improves detection precision and reliability over single-signal or questionnaire-based methods. If that is right, continuous, objective fatigue assessment could be embedded in workplace safety, transportation, sports, and healthcare. The review also maps the field's main barriers: no standard definition of fatigue, predominantly small lab-based studies, data quality and privacy concerns, and the need for explainable and uncertainty-aware models.","feed_headline":"AI wearables can track fatigue in real time","feed_subtitle":"Fusing multiple physiological signals, AI wearables detect fatigue accurately in real time, says a 150-study review.","key_machinery":"The central machinery is the physiological-signal fusion pipeline: wearables continuously capture signals such as heart-rate variability, brain electrical activity, muscle electrical activity, and eye movement; signal processing converts raw traces into features; and machine-learning or deep-learning classifiers map those features to fatigue levels. The key mechanism is information fusion—combining complementary modalities—which the review identifies as the strongest route to accurate, real-time assessment. The paper's tables document this machinery across studies, listing modality, classifier, sample size, experimental setting, and reported accuracy.","core_discovery":"The central claim is that wearable technologies combined with AI, and especially with multi-source data fusion, have substantially improved the precision, real-time capability, and efficiency of fatigue detection. The authors support this by synthesizing around 150 studies that extract features from physiological signals such as ECG, EEG, EMG, PPG, EOG, EDA, and IMU and classify fatigue with machine-learning and deep-learning models, with reported accuracies frequently above 90 percent. They observe that hybrid multimodal models are the most prevalent approach and that SVM is the most popular single classifier, while deep learning is gaining ground. The review also acknowledges that lab and simulation studies dominate, and that real-world deployment still faces hurdles in real-time data access, device ergonomics, data quality, and model transparency.","pith_inferences":["Because most reported accuracies come from small lab or simulation samples, the review's optimistic conclusion likely overstates readiness for real-world deployment; external validation on diverse free-living populations is the next test.","Since ground truth in many studies is self-reported or task-induced, high accuracy may reflect detection of sleepiness, stress, or low arousal rather than fatigue per se; separating these will require benchmark tasks that label fatigue independently of its confounds.","The recurring superiority of hybrid over single-modality models suggests a general design principle: fatigue is multidimensional, so any single biomarker eventually saturates, and fusion of orthogonal signals is needed for robust monitoring.","If edge computing matures as the paper expects, the bottleneck shifts from sensor accuracy to label quality; a standardized, objectively anchored fatigue scale would likely do more for generalizable models than adding more sensors."],"forward_implications":["Workplaces such as transport, construction, aviation, and mining could move from periodic subjective fatigue surveys to continuous, unobtrusive monitoring that alerts before performance drops.","Consumer wearables such as smartwatches and smart glasses could carry validated fatigue models for self-management, provided the models are validated outside the lab.","The identified gaps point to near-term research priorities: standardized fatigue labels, larger field studies, explainable models for trust, and uncertainty quantification for high-stakes decisions.","Edge computing could enable on-device processing, cutting latency and privacy risks for real-time alerts.","Combining physiological signals with behavioral data such as facial video or head motion appears to push accuracy higher, suggesting hybrid sensing as the field's likely direction."],"supporting_citations":[{"why":"Hybrid EEG and EOG signals with a fast SVM for driver fatigue recognition, supporting multi-signal fusion claims.","marker":"[28]"},{"why":"ECG-based driver state classification with multiple classifiers, grounding ECG-based fatigue detection.","marker":"[43]"},{"why":"ECG feature classification for mental fatigue with random forest, showing high accuracy from a single cardiac signal.","marker":"[84]"},{"why":"EEG dry-headband passive fatigue quantification with negative-unlabeled learning, supporting EEG-based real-time monitoring.","marker":"[99]"},{"why":"Multimodal wearable sensors for physical and cognitive fatigue with random forest and LSTM, underpinning the fusion advantage.","marker":"[130]"},{"why":"ECG and video multimodal fusion for learning fatigue detection, a load-bearing example of hybrid methods.","marker":"[82]"},{"why":"Deep learning IoT multimodal fatigue monitoring with eye and mouth state classification, illustrating hybrid and high-accuracy results.","marker":"[71]"},{"why":"Wearable device system monitoring stress, fatigue, and drowsiness, grounding real-world applicability claims.","marker":"[74]"}],"fun_headline_variants":["AI wearables read fatigue signals in real time","150 studies show AI wearables spot fatigue","Wearables and AI: real-time fatigue detection on the rise","Review: AI + wearables see fatigue coming","Fusing signals, AI wearables track fatigue live"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review assumes that the fatigue labels used in the collected studies—often self-reports or task-induced drowsiness in groups of 5 to 64 volunteers—actually measure fatigue rather than correlated states such as sleepiness, stress, or low arousal.","fun_headline_variants_meta":{"raw":{"variants":["AI wearables read fatigue signals in real time","150 studies show AI wearables spot fatigue","Wearables and AI: real-time fatigue detection on the rise","Review: AI + wearables see fatigue coming","Fusing signals, AI wearables track fatigue live"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000221,"raw_usage":{"total_tokens":1452,"prompt_tokens":948,"completion_tokens":504,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":429}},"tokens_in":564,"tokens_out":504,"duration_ms":4380,"temperature":1.0,"reasoning_tokens":429,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:14:52.358402+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A field study equipping several hundred shift workers with multisignal wearables and probing fatigue with an objective benchmark, such as psychomotor vigilance task reaction times at random times, could settle whether the reported lab accuracies generalize. If the AI models' fatigue predictions match the probes better than chance, the central claim holds; if not, the review's conclusion would need to be weakened.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Hybrid EEG and EOG signals with a fast SVM for driver fatigue recognition, supporting multi-signal fusion claims."},{"cited_title":"Murugan, J","cited_arxiv_id":null,"evidence_quote":"ECG-based driver state classification with multiple classifiers, grounding ECG-based fatigue detection."},{"cited_title":"Foong, K","cited_arxiv_id":null,"evidence_quote":"EEG dry-headband passive fatigue quantification with negative-unlabeled learning, supporting EEG-based real-time monitoring."},{"cited_title":"Assessing Fatigue with Multimodal Wearable Sensors and Machine Learning","cited_arxiv_id":"2205.00287","evidence_quote":"Multimodal wearable sensors for physical and cognitive fatigue with random forest and LSTM, underpinning the fusion advantage."}],"review_version":1}