{"id":"f0dd3f5c-645b-48fe-8b12-1cecc70764a5","arxiv_id":"2501.04156","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A neuroadaptive LLM cockpit assistant increased fNIRS-labeled optimal working-memory states in eight pilots, but the effect was measured with the same classifier the system is built to please.","lead":"AdaptiveCoPilot is a virtual reality cockpit guidance system that reads a pilot's brain activity and changes its voice, visual, and text instructions in real time to try to keep the pilot's mental workload in a good range. In a small test with licensed pilots it did raise the rate of the system's own 'optimal' working-memory classification, but the tiny sample and the circular success measure limit what can be concluded.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The primary outcome is the same classifier output that the adaptive controller optimizes; without anchoring 'optimal' labels to objective performance, the significant memory-state result may be a feedback artifact rather than evidence of improved cognitive state.","rationale":"The paper is best read as a proof-of-concept, and the authors are unusually candid in Secs. 7 and 8: they acknowledge the memory-only result, the LLM's confusion on attention/perception, and the small sample. My concern does not allege artifact in the statistical models; it targets the construct validity of the primary outcome. AdaptiveCoPilot is a closed-loop system: the PHI-3 policy consumes the classifier's three-facet labels and selects modalities and information load to push those labels toward 'optimal', and the evaluation measures the incidence of those same labels. A classifier can be well-calibrated on its training distribution and still be an invalid optimization target if the policy changes features that the classifier uses as shortcuts, or if the 'optimal' class does not correspond to better performance. The available objective measures do not help: completion time is not significant, and the error analysis shows the guidance conditions incurred more errors, not fewer. The perception result is further confounded because the random condition produced more optimal perception labels than the adaptive condition, suggesting the label shift tracks modality presentation rather than adaptive intelligence. The requested check uses existing baseline data to test whether optimal memory labels predict objective performance within the no-guidance condition; if they do not, the significant label-based result is not interpretable as an operational improvement. I therefore keep the reader's CONDITIONAL verdict: the paper is a worthwhile case study, but the central quantitative claim needs this validation before acceptance as evidence of neuroadaptive benefit.","tokens_in":19329,"tokens_out":5494,"duration_ms":56600,"concrete_test":"Using the recorded baseline-condition ROS BAGs, fit a mixed-effects model of objective step times and error counts on concurrent memory-facet classifier labels (optimal vs non-optimal), within the baseline condition where no adaptive feedback is present. If 'optimal' memory labels do not predict faster steps or fewer errors (or if the association reverses), the classifier lacks construct validity as a workload/performance measure in this task, and the significant adaptive-vs-baseline label difference cannot support the claim that AdaptiveCoPilot improves cognitive state. As a secondary check, repeat the same analysis in the random condition to test whether non-adaptive modality changes produce similar label shifts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"AdaptiveCoPilot's central quantitative claim (Sec 6.3, Fig. 5) is that adaptive guidance increases rates of 'optimal' cognitive-load labels. But those labels are produced by the same multinomial regression classifiers (Sec 5.2.1, from ref. [42]) whose outputs drive the PHI-3 adaptation policy (Sec 5.2.2). The system is explicitly designed to move the pilot into the classifier's 'optimal' region; the evaluation then counts how often the pilot is in that region. This closed loop is only meaningful if the labels are causally tied to actual workload and performance in this VR preflight task. The paper provides no such validation: completion time differences are not significant (Sec 6.3: p=0.3034 and p=0.3998 for adaptive vs baseline and random), and the error analysis shows a significant difference between adaptive and baseline (p=0.0228), with the Discussion describing an increase in error counts in the guidance conditions. The perception result is also confounded: the random condition produced more optimal perception labels than the adaptive condition (random-vs-adaptive beta=0.421 versus baseline-vs-adaptive beta=-1.403), suggesting label shifts track modality presentation rather than adaptive intelligence. The authors themselves note the LLM was confused about attention/perception distinctions (Sec. 7), and two participants' fNIRS sessions were unusable (Sec. 6.1). If the classifier's 'optimal' label is not anchored to performance, the significant memory-label advantage could reflect modality-induced shifts in fNIRS features rather than improved cognitive state.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"AdaptiveCoPilot is a VR preflight guidance system that uses fNIRS-based classifiers of working memory, perception, and attention to drive a PHI-3 LLM that adapts the modality and information load of cockpit instructions. A formative interview study with three expert pilots produced design requirements and 28 adaptive rules, and an eight-pilot within-subject VR experiment compared baseline (paper checklist), random guidance, and adaptive guidance. The paper reports a statistically significant increase in 'optimal' working-memory label rates under adaptive guidance relative to both baseline and random conditions, and it reports non-significant completion-time differences plus a significant difference in error counts whose direction is stated inconsistently. The Discussion honestly narrows the central claim to the working-memory classifier result, but the Abstract and Conclusion overstate the findings by presenting reduced completion time and improved perception states as established results.","tokens_in":19628,"tokens_out":6091,"duration_ms":63879,"significance":"If the fNIRS classifier labels are valid for this VR preflight task, the working-memory result is an interesting proof-of-concept for closed-loop neuroadaptive training guidance, and the paper's qualitative findings plus design strategies are useful for future adaptive cockpit systems. The manuscript also benefits from an unusually candid Discussion that acknowledges the perception, attention, completion-time, and error findings did not support the original hypotheses. However, the primary quantitative outcome is the rate of 'optimal' labels produced by the same classifiers whose outputs drive the adaptive controller, and no independent validation anchors those labels to objective task performance in this study; the paper's significance therefore hinges on an inherited and unverified classifier, the underlying model of which is not released.","major_comments":[{"comment":"The central quantitative claim is circular in structure. The adaptive policy in §4.1 and §5.2.2 is explicitly designed to move the fNIRS classifier outputs into the 'optimal' region, and the primary success measure in §6.3 is the rate of 'optimal' labels from the same classifiers introduced in §5.2.1. Without an independent validation connecting these labels to objective performance (e.g., step accuracy, error recovery, or task time within this VR preflight task), the significant working-memory result may be a feedback artifact rather than evidence of improved cognitive state. The paper should provide such an anchoring analysis, or explicitly reframe the contribution as 'increases in the classifier's optimal-label rate' with the circularity stated as a limitation.","section":"§5.2.1, §5.2.2, §6.3"},{"comment":"The Abstract's claim that AdaptiveCoPilot produced 'higher rates of optimal cognitive load states on the facets of working memory and perception' is not supported by the paper's own statistics. In §6.3, the random-vs-adaptive comparison for perception has beta = 0.421 (p < 0.001), meaning the random condition had a higher optimal-perception label rate than the adaptive condition; only the baseline-vs-adaptive comparison favored adaptive. The Discussion (Sec. 7) correctly acknowledges that 'the random condition was better then both,' so the Abstract and the 'Results indicate' sentence should be corrected.","section":"Abstract and §6.3 Perception analysis"},{"comment":"The Abstract and Conclusion state that AdaptiveCoPilot 'accelerated task completion time' and 'reduced task completion times,' but the completion-time gamma model in §6.3 found no significant differences (baseline vs adaptive p = 0.3034; random vs adaptive p = 0.3998). A non-significant trend toward faster completion is not the same as an accelerated completion time. The Conclusion's sentence 'Results show that AdaptiveCoPilot accelerated task completion time relative to the baseline and random system condition' is therefore unsupported and must be revised, especially in light of the Limitations section's own warning that quantitative findings should be read as indicative trends.","section":"Abstract, §9 Conclusion, §6.3 Completion Time"},{"comment":"The error-count results are internally inconsistent. The text reports a rate ratio of 0.644 for adaptive vs baseline (p = 0.0228), which, under the stated parameterization, means the adaptive condition had roughly 35% fewer errors than baseline; yet the same paragraph concludes 'a lower error count with baseline relative to adaptive,' and the Discussion (§7) and strategy list (§7.1) describe 'an increase in error counts in the guidance conditions' and 'higher error rates compared to the baseline.' The manuscript must state the model's reference coding unambiguously and reconcile these conflicting descriptions, because the direction of the error effect is load-bearing for the paper's interpretation of complacency and speed-accuracy trade-offs.","section":"§6.3 Error Counts"},{"comment":"The mixed-effects models are not specified in enough detail to assess the reported p-values. The working-memory model yields z = -16.173 and z = -30.737 with only eight participants (and, per §6.1, only six usable fNIRS sessions), which suggests that the unit of analysis may be individual 10 Hz time windows or per-procedure observations rather than participants. The manuscript should report the exact model formulas, the random-effects structure, the definition of the observation unit, and the effective sample size used in each analysis. Without this information, the extremely small p-values for the memory result cannot be trusted, and a pseudoreplication risk remains.","section":"§6.3 Statistical model specification"},{"comment":"The manuscript reports that fNIRS sessions from two of the eight quantitative participants 'could not be used in our final quantitative evaluation,' but the subsequent statistical analyses in §6.3 do not state that all fNIRS-based results are based on N = 6 rather than N = 8. The sample size for each model should be reported explicitly, and the implications for statistical power and generalizability should be discussed, especially since the Limitations section emphasizes the small sample.","section":"§6.1 and §6.3"}],"minor_comments":[{"comment":"Several typos should be corrected: 'Prepossessing' for 'Preprocessing' (§5.1), 'cognivive' (Introduction), 'Similarily' (§2), 'neuoradaptive' (§1), and 'there pilots' (§7.1).","section":"Throughout"},{"comment":"The supplemental list numbers two items as '(5)' (qualitative evaluation questions and prompt examples); the numbering should be fixed, and the file names should be listed explicitly so readers can locate each artifact.","section":"§10 Supplemental Materials"},{"comment":"The sentence 'Our classifiers rely on multinomial symbolic regression, previously trained on a Rasch model labeling methodology, updating at 10hz' is ambiguous about whether the classifier updates at 10 Hz or the input features do; please rephrase and specify the temporal granularity of the labels.","section":"§5.2.1"},{"comment":"The figures show raw trends across procedures and conditions but do not include confidence intervals or model-based estimates; adding error bars or shaded intervals would make the reported mixed-effects contrasts easier to interpret.","section":"Figures 4–7"},{"comment":"References [3] and [4] appear to describe the same work with different bibliographic metadata; please verify and retain only the correct source.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The circularity concern is real and should be central to the revision: the adaptive controller optimizes the same classifier labels that are then counted as the outcome. I would not require the authors to abandon the working-memory finding, but I would ask them to either re-analyze existing data to anchor the labels to objective performance (errors, time, action correctness) or explicitly reposition the paper as a proof-of-concept about classifier-label control. Given that the classifier is co-authored by a co-author of this manuscript and the trained models are proprietary, independent validation is especially important. The inconsistent error-rate direction and the unsupported completion-time claim in the Abstract/Conclusion are fixable in revision and should be addressed before a decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real system, tested with real licensed pilots, and the working-memory effect is statistically significant. But the abstract claims things the data don't support (completion time, perception), and the central metric is the fNIRS classifier's own 'optimal' label, which the adaptive controller is explicitly trying to maximize. That circularity doesn't kill the paper as a proof-of-concept, but it should be front and center.\n\nWhat's new: the integration. fNIRS workload classification, LLM reasoning, and multimodal guidance in a VR preflight checklist—I haven't seen that combination before. The formative study with three experts is a reasonable way to ground the adaptive rules, and the qualitative interviews add color. The authors are also honest in the Discussion: they admit the memory result is the only strong one, that the LLM was confused about attention vs perception, and that error rates went up in the guidance conditions.\n\nSoft spots, in order of softness. First, the primary outcome is the classifier output itself. The system reads fNIRS, classifies into underload/optimal/overload, and then picks guidance to push the pilot toward 'optimal'. The evaluation then counts how often the pilot is 'optimal'. If the classifier is valid, this is meaningful; if not, it's a feedback loop. The classifier comes from a co-author's prior work and isn't independently validated in this VR preflight task. That's a load-bearing assumption.\n\nSecond, the abstract and conclusion overreach. Completion time is not significant (p=0.30, 0.40), and the error analysis shows baseline has significantly fewer errors than adaptive (p=0.023). The perception result actually favors random over adaptive. So the paper's headline claims don't match the paper's own numbers.\n\nThird, the small sample: eight pilots, with two fNIRS sessions unusable, so the quantitative analysis rests on six. The repeated-measures design and mixed-effects models help, but this is still a case study.\n\nNone of this makes the paper worthless. The memory optimal-state effect is in the predicted direction and highly significant (p<0.001 for both comparisons), and the system is a credible prototype. What's needed is a larger, preregistered study that anchors the workload labels to objective performance, or at least a clear statement that the current result is about classifier states, not real-world outcomes.\n\nBottom line: this deserves a serious referee. The idea is timely, the system is real, and the authors are transparent about limitations. I'd recommend major revision with a focus on reframing the claims and adding a performance-based validation of the classifier. For a reading group, it would generate good discussion about how to evaluate neuroadaptive systems.\n\nMy call: accept for peer review, cite cautiously.","headline":"A genuine proof-of-concept for LLM-driven neuroadaptive guidance, but the abstract oversells the results and the main outcome is the same classifier the system is designed to optimize.","tokens_in":20262,"tokens_out":3231,"would_cite":false,"duration_ms":30658,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AdaptiveCoPilot claims that a closed-loop system reading fNIRS workload signals and using a small language model to adapt cue modality and information load keeps pilots in optimal working-memory states during VR preflight checklists, with…","keywords":["neuroadaptive system","cognitive workload","fNIRS","large language model","virtual reality","pilot training","adaptive guidance","preflight checklist"],"falsifier":"A concrete test: compute, within each 10-second window, whether the proportion of time the memory classifier spends in 'optimal' predicts the probability of a procedure error or a step-completion delay in the next window. The paper already contains a hint in this direction, because error counts rose under adaptive guidance while optimal memory labels also rose; if that inverse relationship holds in the recorded ROS data, the 'optimal' label is not a valid proxy for performance. A second, cleaner experiment would run the same LLM guidance with the classifier labels replaced by labels generated at the same marginal rates but uncorrelated with behavior, which would isolate whether the adaptation is responding to the brain signal or merely to the guidance schedule.","tokens_in":19107,"feed_emoji":"🧠","tokens_out":9603,"duration_ms":83853,"temperature":0.7,"pith_summary":"The paper claims that a cockpit assistant can read a pilot's cognitive workload from brain signals and adapt its own guidance in real time, and that this keeps pilots in a healthier 'optimal' workload state—at least for working memory—during a VR preflight checklist. The system, AdaptiveCoPilot, closes a loop: fNIRS (functional near-infrared spectroscopy) prefrontal readings are classified into underload, optical, and overload states for memory, perception, and attention; a small language model then chooses which cue modality (visual, audio, text) and how much information to deliver. The authors report a significant increase in optimal working-memory states for the adaptive condition against both a paper-checklist baseline and a randomly guided control, and treat this as a proof-of-concept that neuroadaptive LLM guidance is feasible. They are explicit that the broader hypothesis—better workload leading to faster, error-free performance—was only partially supported, and they frame their results as indicative trends from a small pilot sample.","feed_headline":"Brain-reading cockpit assistant boosts optimal memory states","feed_subtitle":"Adapting cues to live brain-workload signals beat checklist and random guidance on working memory in VR preflight tests.","key_machinery":"The load-bearing mechanism is the closed loop among three components: the fNIRS preprocessing and classification pipeline (wavelet filtering, a sliding 10-second window, and three separate multinomial classifiers for working memory, perception, and attention, each trained with a Rasch-model labeling methodology from prior work), a set of 28 adaptive rules derived from expert interviews and workload theory that map underload, optimal, and overload states to modality and information-density choices, and a quantized PHI-3 LLM prompted with chain-of-thought reasoning that turns the current state, task tree, and gaze into a concrete guidance action. The claimed effect is that this loop produces more 'optimal' labels on the memory classifier than fixed or random guidance does.","core_discovery":"On its own terms, the central discovery is that an adaptive guidance loop driven by a real-time workload classifier can move the classified state itself: pilots using AdaptiveCoPilot had significantly higher rates of 'optimal' labels on the working-memory classifier than pilots using either the checklist baseline or a same-cadence random guidance condition, a difference the authors describe as strong evidence. The same effect did not hold for attention, and for perception the random condition outperformed the adaptive one even though both beat the baseline. Completion time trended lower with adaptive guidance but was not statistically significant, and error counts were significantly higher in the adaptive condition than in the baseline, which the authors attribute to complacency and a possible speed-accuracy trade-off. The paper's conclusion is therefore narrower than its framing: the system's demonstrated effect is on working-memory workload classification, and the authors recommend future work on the perceptual and attentional rules, on complacency, and on the LLM's evident confusion between attention and perception states.","pith_inferences":["The strongest experimental contrast in the paper is adaptive vs random at a fixed 10-second cadence; a stricter test of 'neuroadaptivity' would compare against guidance selected by the same LLM from the same context but with the classifier labels withheld, isolating whether the brain-derived signal is what drives the effect or whether the LLM's contextual reasoning alone would do as well.","The classifier labels come from models trained on a Rasch-based labeling method that defines 'optimal' as a midpoint of capability; driving pilots toward that midpoint may be beneficial mainly for trainees who are below their capacity, and the same rule could degrade experts who are already near their ceiling—the paper's own expertise interviews hint at this.","One could extend the study to other high-stakes procedural work (medical checklists, ATC handovers) since the mechanism—fNIRS workload classification feeding an LLM that adapts modality and detail—is task-agnostic; the paper does not claim this, but the design generalizes.","Because the paper found no significant completion-time benefit and a significant error increase, the 'optimal' label may not be aligned with objective performance; a direct test would be whether the proportion of optimal-classified time in a window correlates with next-step error probability."],"forward_implications":["If the working-memory effect replicates, real-time fNIRS-based adaptation could be built into cockpit training aids, not just for UH-60 preflight but for any checklist-driven procedure where working memory is the bottleneck.","The result implies that adaptive rules do affect the classified state: the adaptive condition beat both a no-guidance baseline and a same-rate random guidance condition on memory, so the adaptation itself—not merely the presence of feedback—is what moves the classifier.","Because the random condition improved perception more than the adaptive one, the paper's own conclusion is that its strategies for managing perceptual and attentional load are ineffective; future systems need separate, better-tuned rules for those facets.","The increased error rates under adaptive guidance suggest a complacency or speed-accuracy trade-off; the authors recommend future systems explicitly manage complacency, for example by flagging when a pilot is taking too long on a step.","The qualitative interviews point to training, rather than operational flight, as the most plausible near-term deployment: experts said experienced pilots would not want the system during routine operations, but novices could benefit."],"supporting_citations":[{"why":"Supplies the fNIRS preprocessing, the three cognitive-facet classifiers, and the Rasch-based label definitions of underload, optimal, and overload that the adaptive loop uses as its live inputs.","marker":"[42]"},{"why":"Establishes the curvilinear relationship between working-memory load and prefrontal hemodynamics, motivating the targeting of an optimal midpoint rather than simply minimizing workload.","marker":"[43]"},{"why":"Provides the Yerkes-Dodson arousal-performance curve used as the theoretical justification for defining an optimal cognitive-load state.","marker":"[17]"},{"why":"Wickens' Multiple Resource Theory is the basis for the modality-switching strategy: distributing information across visual, auditory, and textual channels to manage limited cognitive resources.","marker":"[57]"},{"why":"Documents the human-factors role of flight-deck checklists and provides the procedural context (the paper checklist) that the system augments and against which it is compared.","marker":"[18]"},{"why":"Argues that discrete state transitions matter more than fine-grained workload fluctuations and that subjective measures are biased, supporting the design decision to classify into underload, optimal, and overload states.","marker":"[44]"}],"fun_headline_variants":["Adaptive brain-sensing copilot improves working-memory states","Neuroadaptive LLM copilot: working-memory gain, but not attention","Cockpit AI that reads brain load helps memory, not perception","Adaptive brain-guided copilot boosts memory, but more errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole outcome measure rests on the pre-trained fNIRS classifiers correctly labeling each moment as underload, optimal, or overload for each cognitive facet in this specific VR preflight task; if those labels are miscalibrated here, then increasing the frequency of 'optimal' labels may not correspond to any real improvement in pilot workload or performance.","fun_headline_variants_meta":{"raw":{"variants":["Adaptive brain-sensing copilot improves working-memory states","Neuroadaptive LLM copilot: working-memory gain, but not attention","Cockpit AI that reads brain load helps memory, not perception","Adaptive brain-guided copilot boosts memory, but more errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001399,"raw_usage":{"total_tokens":5671,"prompt_tokens":975,"completion_tokens":4696,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":4621}},"tokens_in":591,"tokens_out":4696,"duration_ms":34832,"temperature":1.0,"reasoning_tokens":4621,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:39:36.996861+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: compute, within each 10-second window, whether the proportion of time the memory classifier spends in 'optimal' predicts the probability of a procedure error or a step-completion delay in the next window. The paper already contains a hint in this direction, because error counts rose under adaptive guidance while optimal memory labels also rose; if that inverse relationship holds in the recorded ROS data, the 'optimal' label is not a valid proxy for performance. A second, cleaner experiment would run the same LLM guidance with the classifier labels replaced by labels generated at the same marginal rates but uncorrelated with behavior, which would isolate whether the adaptation is responding to the brain signal or merely to the guidance schedule.","supporting_citations":[{"cited_title":"McKendrick, B","cited_arxiv_id":null,"evidence_quote":"Supplies the fNIRS preprocessing, the three cognitive-facet classifiers, and the Rasch-based label definitions of underload, optimal, and overload that the adaptive loop uses as its live inputs."},{"cited_title":"McKendrick and A","cited_arxiv_id":null,"evidence_quote":"Establishes the curvilinear relationship between working-memory load and prefrontal hemodynamics, motivating the targeting of an optimal midpoint rather than simply minimizing workload."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Yerkes-Dodson arousal-performance curve used as the theoretical justification for defining an optimal cognitive-load state."},{"cited_title":"Degani and E","cited_arxiv_id":null,"evidence_quote":"Documents the human-factors role of flight-deck checklists and provides the procedural context (the paper checklist) that the system augments and against which it is compared."}],"review_version":1}