{"id":"ae30d5c5-6e17-4a28-9725-9ac5e0108e87","arxiv_id":"2508.05025","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Eye-tracking graphs from an AR CPR task predict freeze-probe situational awareness scores with 83% accuracy; larger, faster saccades and less time on virtual content mark higher awareness.","lead":"People doing CPR in augmented reality can stare at virtual guidance and miss real emergencies. The authors tracked eye movements while simulated incidents occurred, then trained a graph neural network that predicted situational awareness from gaze with 83% accuracy; more aware users made larger, faster saccades and spent less time on virtual content.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"83% accuracy may reflect incident-identity leakage rather than general SA: freeze-probe labels and gaze features share the same noticing event, and the three incidents differ sharply in modality and location.","rationale":"The reader identified the same vulnerability: the freeze-probe label and the gaze features share a causal source, and the incidents differ in modality and detectability. My reading of the available sections confirms this is the main point that must be settled before the 83% claim can be taken as evidence for gaze-based SA modeling in general. The internal design is careful in other ways—the 21 cm placement for blood and vomit is an explicit attempt to equate detection difficulty, a pilot study informed the app, and a CPR instructor validated incidents—but this does not control the ambulance's completely different visual/auditory signature. The missing Sections 4-7 are decisive only in that they prevent checking whether the authors already addressed this via within-incident splits or label construction; therefore the correct disposition is conditional acceptance pending the incident-identity check, not rejection. I therefore leave the reader's verdict unchanged.","tokens_in":10017,"tokens_out":4620,"duration_ms":55277,"concrete_test":"Run leave-one-incident-out cross-validation on the same gaze-event graphs used for the 83% result: train FixGraphPool on blood and vomit trials and test on ambulance trials, then cycle the held-out incident. Compare against chance and against a trivial baseline that uses only a binary 'fixation entered the incident AOI before the freeze probe' feature. As a second check, train the model to predict incident identity (B/V/A) from the same feature set; if incident-identity accuracy is near 100% while held-out-incident SA accuracy drops to near chance, the SA results are confounded by incident-specific gaze signatures. If instead accuracy is stable across held-out incidents and substantially beats the fixation-baseline, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that FixGraphPool classifies situational awareness from eye tracking at 83% accuracy. The load-bearing requirement is that the label is not already contained in the gaze features through a trivial perceptual channel. The freeze-probe SA label is scored substantially from whether the participant noticed a staged incident; noticing is directly manifest as fixations on the incident and saccades away from the AR overlay. The gaze features (fixation proportion/frequency on virtual content, saccadic amplitude/velocity) will therefore be informative for the label by construction. The three incidents are not symmetric: blood and vomit are physical releases 21 cm from the compression point, while the ambulance is a virtual object approaching over 95 m with a siren, a different spatial scale and modality. If the model is trained on gaze windows aligned to incident onset and the label is per-incident, it may learn to recognize which incident is present (or whether the participant happened to look at it) rather than a general SA state. Sections 4-7, which would report the label protocol, feature windows, and split/validation scheme, are absent from the provided text; Sec. 3.2 points to a Sec. 7 limitation discussion that is also missing. Without a test that holds incident identity fixed, the 83% figure does not yet distinguish 'ET can detect noticing/specific incident' from 'ET models situational awareness as a general state.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes to model situational awareness (SA) during AR-guided CPR using eye-tracking data from a Magic Leap 2 headset. It describes the design of an AR CPR guidance app, a user study with three staged incidents (bleeding, vomiting, and a virtual ambulance), and a graph neural network model, FixGraphPool, that encodes gaze events as spatiotemporal graphs. The abstract reports 83.0% accuracy (F1=81.0%) for SA classification and states that higher SA is associated with larger saccadic amplitude/velocity and lower fixation proportion/frequency on virtual content. However, the provided manuscript does not include the evaluation, model, or results sections (Sections 4–7), so the central empirical claims cannot be audited from the text supplied.","tokens_in":10234,"tokens_out":2501,"duration_ms":31058,"significance":"If the claims hold, this would be a valuable demonstration that eye-tracking from a commercial AR headset can classify an operator's situational awareness in a dynamic safety-critical task, with direct implications for gaze-based cognitive-tunneling detection in AR guidance. The task design in Section 3 is thoughtful: incidents were selected with a certified CPR instructor, a pilot study informed the app, and the bleeding/vomiting releases were positioned to equate detection difficulty. The proposed FixGraphPool architecture is a reasonable domain-informed approach. However, the validity of the 83% accuracy figure depends on whether the SA labels are independent of the gaze features and whether the model generalizes across incident types, both of which are unresolved in the provided text.","major_comments":[{"comment":"The core empirical claim—83.0% accuracy (F1=81.0%)—is stated in the abstract, but the provided manuscript contains no evaluation section, no model description beyond the name FixGraphPool, no feature-window definition, and no train/validation split details. Section 3.2 even references a Section 7 limitations discussion that is absent. Without these components, the headline accuracy cannot be verified, and baseline comparisons cannot be assessed. This is a load-bearing omission that must be remedied before the claims can be evaluated.","section":"Abstract / Sections 4–7"},{"comment":"The SA labels appear to be assigned substantially from whether the participant noticed the staged incidents (freeze-probe observation/questionnaires). Noticing is directly manifest in eye movements: fixations on the blood/vomit/ambulance and saccades away from the AR overlay. Since FixGraphPool encodes exactly these gaze events, the model may be reconstructing the labeling process rather than predicting an independent SA construct. The paper should report the exact SA scoring rubric, separate label collection from gaze features, and demonstrate with a control analysis (e.g., hold out all windows where the incident is fixated) that accuracy does not collapse.","section":"Section 3.3 / label-feature circularity"},{"comment":"The three incidents are not perceptually symmetric. Bleeding and vomiting are physical releases positioned 21 cm from the compression point, while the ambulance is a virtual object approaching from 95 m at 6 m/s with an audible siren—a different modality, spatial scale, and salience profile. If the model is trained on per-incident windows, it can learn to identify which incident is present (or whether the participant's gaze happened to land on the incident) rather than a general SA state. A leave-one-incident-out evaluation, or at least reporting accuracy per incident, is necessary to support the general claim of SA classification.","section":"Section 3.3 / incident comparability"},{"comment":"The freeze-probe SA measurement procedure is not described in the provided sections. The paper mentions 'observation and questionnaires administered during freeze-probe events' but gives no details on the number of probes, their timing relative to incident onset, or the questionnaire items. Since the entire label construction depends on this protocol, its absence is a critical gap that prevents assessment of label reliability and of potential leakage from the labeling procedure into the gaze features.","section":"Section 3.2 / freeze-probe protocol"}],"minor_comments":[{"comment":"The title contains odd spacing: 'Will Y ou Be Aware?' should be 'Will You Be Aware?'.","section":"Title/header"},{"comment":"The DOI placeholder 'xx.xxxx/TVCG.201x.xxxxxxx' is unresolved and should be completed or removed before publication.","section":"Copyright/front matter"},{"comment":"The text refers to 'as illustrated in Fig. 2', but Figure 2 is not present in the provided text. Please verify the figure is included and legible.","section":"Section 3.3 / Figure 2"},{"comment":"The description of the ambulance incident states it 'would approach with a siren sound from the front-left' but does not specify whether the siren is spatialized or how its volume interacts with the CPR drumbeat at 108 BPM. This could affect detection difficulty and should be reported.","section":"Section 3.3 / incident implementation"}],"recommendation":"major_revision","confidential_remarks":"The provided manuscript is missing the entire evaluation and model sections (Sections 4–7), despite the abstract presenting strong empirical results. It is possible that the version sent for review was truncated; if so, I recommend requesting the complete manuscript before further review. The core idea is promising, but the circularity and incident-asymmetry concerns need to be addressed with explicit control analyses."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Josh — quick read of arXiv:2508.05025. The abstract promises 83% accuracy classifying situational awareness from eye tracking during AR-guided CPR. The part I can see looks good: the app design is detailed, the incidents are physically simulated with controlled placement (21 cm from compression point), a pilot study informed the design, and a certified CPR instructor validated the incidents. That is real work, and the reported associations — higher SA with larger saccadic amplitude/velocity and fewer fixations on virtual content — are plausible, if not shocking.\n\nThe soft spots line up with the preliminary review. First, the freeze-probe SA label is partly defined by whether the participant noticed the staged incident. That noticing is directly visible in the gaze stream as a fixation on the incident and/or a saccade away from the overlay. The classifier may be reconstructing the label rather than predicting an independent construct. That's not automatically fatal — if the goal is a practical detector of overt attention, fine — but it does undercut the stronger claim that eye tracking 'models situational awareness as a general state.'\n\nSecond, the three incidents are not symmetric: blood and vomit are physical events at the mannequin; the ambulance is a virtual object approaching from 95 m with a siren. If the model is trained on gaze windows aligned to incident onset, it has an easy cue for which incident is present. Without a held-out analysis that controls for incident identity, the 83% figure could just mean 'ET recognizes the incident type.'\n\nThe real issue: the provided text cuts off around Section 3.3, so Sections 4-7 (procedure, eye-tracking pipeline, FixGraphPool, results, discussion) are missing. We cannot audit the actual evaluation. That's a limitation of the review packet, not necessarily the paper. But it means the central number is unverified from this excerpt.\n\nI wouldn't desk-reject this. The topic is timely, the study design is above average for the area, and the authors are clearly thinking about the construct. I'd send it to a referee with explicit instructions to check whether incident identity and participant identity are accounted for, and to require a statement on the label-feature overlap. If the full paper addresses those, this is a reasonable contribution to the AR/SA literature. If not, the 83% is overstated.\n\nBring it up at reading group if you want a debate on what freeze-probe SA actually measures.","headline":"A careful AR-CPR study with a plausible gaze–SA link, but the headline accuracy can't be audited in the provided text, and the freeze-probe labels may be encoding the very gaze behavior the model sees.","tokens_in":10834,"tokens_out":2821,"would_cite":false,"duration_ms":31622,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Eye-tracking data from a commercial AR headset predicts a user's situational awareness level with 83% accuracy, enabling gaze-based detection of cognitive tunneling.","keywords":["eye tracking","situational awareness","augmented reality","graph neural network","cognitive tunneling","gaze analysis","CPR guidance","spatiotemporal graphs"],"falsifier":"Run the trained model on gaze windows from which the time segments around incident onset have been removed. If accuracy drops to near chance, the model is mainly reading incident detection rather than a general SA state. Alternatively, have new participants perform the same CPR task with the visual overlay but no staged incidents, and check whether the model still separates participants who later report high vs. low awareness.","tokens_in":9786,"feed_emoji":"👁️","tokens_out":4885,"duration_ms":50604,"temperature":0.7,"pith_summary":"The paper tries to establish that eye-tracking signals from a head-mounted AR display contain enough information to infer a user's situational awareness during a dynamic, safety-critical task—specifically AR-guided CPR. Using a graph neural network that represents gaze events as spatiotemporal graphs, the authors report 83.0% accuracy (F1=81.0%) in classifying freeze-probe SA levels, beating feature-based machine learning and state-of-the-art time-series models. The analysis also identifies a concrete gaze signature of high SA: larger and faster saccades and less time spent fixating on virtual content. If this holds, AR guidance systems could continuously monitor attention and detect cognitive tunneling before it leads to missed hazards.","feed_headline":"Gaze data predicts AR users' situational awareness with 83% accuracy","feed_subtitle":"A graph neural network on eye-tracking events detects cognitive tunneling during AR-guided CPR, enabling adaptive safety overlays.","key_machinery":"FixGraphPool, a graph neural network that turns eye-tracking events into spatiotemporal graphs: fixations and saccades become nodes/edges carrying position, duration, and velocity attributes, and a graph-pooling layer aggregates them before classification. The graph structure preserves the order and geometry of gaze, which is what lets the model capture whether the user is scanning the environment or stuck on the virtual content.","core_discovery":"The central discovery is that situational awareness in an AR guidance setting leaves a measurable trace in gaze dynamics, and that this trace can be learned by a graph neural network. In a user study with simulated bleeding, vomiting, and ambulance-arrival incidents during CPR, the authors find that participants with higher SA make saccades with greater amplitude and velocity and spend a smaller proportion of their fixation time on the virtual overlay. They then encode fixations and saccades as nodes and edges of a graph with spatial and temporal attributes, pool them with a GNN, and classify SA level from a short gaze window. The reported 83.0% classification accuracy indicates that the mod","pith_inferences":["The 83% accuracy may partly reflect that SA labels are derived from whether the participant noticed staged incidents, and noticing is itself visible in the gaze stream (fixation on the incident, saccade away from the overlay). A stricter test would mask the gaze segments that coincide with incident onset before classification.","If the saccade-amplitude signature generalizes, it could support a lightweight 'SA index' computed from raw gaze statistics without any neural network, useful for low-power headsets.","The paper's freeze-probe protocol yields discrete SA labels; a continuous variant (e.g., rating awareness at random intervals) would let the same graph architecture output a moment-by-moment SA estimate, which is what an adaptive AR system would actually need."],"forward_implications":["Gaze can serve as a real-time, passive SA sensor in AR: the model runs on the headset's eye-tracking stream and could flag when an operator's awareness drops.","The identified gaze signature—high saccadic amplitude/velocity, low virtual-fixation share—gives interface designers a measurable target for reducing cognitive tunneling.","Graph-based encoding of gaze events outperforms both handcrafted feature classifiers and raw time-series deep learning, suggesting the spatiotemporal structure of gaze carries the predictive signal.","The approach transfers in principle to other safety-critical AR tasks (e.g., first response, remote guidance) where unexpected events demand environmental vigilance."],"supporting_citations":[{"why":"Supplies the definition of situational awareness (perceive, comprehend, project) that motivates the task and the freeze-probe measurement.","marker":"[9]"},{"why":"The SAGAT freeze-probe technique is the basis for collecting the ground-truth SA labels used to train and evaluate the model.","marker":"[15]"},{"why":"GazeGraph establishes the graph-based representation of gaze events for cognitive context sensing, which FixGraphPool extends.","marker":"[22]"},{"why":"The I-VT fixation filter converts raw gaze samples into the fixations and saccades that form the nodes and edges of the graph.","marker":"[62]"},{"why":"TSMixer serves as a state-of-the-art time-series baseline that the proposed model is compared against and outperforms.","marker":"[77]"}],"fun_headline_variants":["Gaze patterns reveal AR users' situational awareness","Eye tracking predicts AR safety awareness with 83% accuracy","Graph neural net maps gaze to detect AR cognitive tunneling","AR gaze analysis spots lapses in situational awareness","Eye-tracking model flags AR-induced tunneling in real time"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The freeze-probe SA labels are treated as a ground truth that is not already contained in the gaze features, but SA is scored largely by whether the participant noticed the incidents, and incident detection is directly visible as a gaze pattern, so the label and the features may share a causal source.","fun_headline_variants_meta":{"raw":{"variants":["Gaze patterns reveal AR users' situational awareness","Eye tracking predicts AR safety awareness with 83% accuracy","Graph neural net maps gaze to detect AR cognitive tunneling","AR gaze analysis spots lapses in situational awareness","Eye-tracking model flags AR-induced tunneling in real time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000261,"raw_usage":{"total_tokens":1448,"prompt_tokens":779,"completion_tokens":669,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":593}},"tokens_in":523,"tokens_out":669,"duration_ms":6709,"temperature":1.0,"reasoning_tokens":593,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:36:13.879720+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained model on gaze windows from which the time segments around incident onset have been removed. If accuracy drops to near chance, the model is mainly reading incident detection rather than a general SA state. Alternatively, have new participants perform the same CPR task with the visual overlay but no staged incidents, and check whether the model still separates participants who later report high vs. low awareness.","supporting_citations":[{"cited_title":"Toward a theory of situation awareness in dynamic systems,","cited_arxiv_id":null,"evidence_quote":"Supplies the definition of situational awareness (perceive, comprehend, project) that motivates the task and the freeze-probe measurement."},{"cited_title":"Situation awareness global assessment technique (SAGAT),","cited_arxiv_id":null,"evidence_quote":"The SAGAT freeze-probe technique is the basis for collecting the ground-truth SA labels used to train and evaluate the model."},{"cited_title":"GazeGraph: Graph- based few-shot cognitive context sensing from human visual behav- ior,","cited_arxiv_id":null,"evidence_quote":"GazeGraph establishes the graph-based representation of gaze events for cognitive context sensing, which FixGraphPool extends."},{"cited_title":"The Tobii I-VT fixation filter,","cited_arxiv_id":null,"evidence_quote":"The I-VT fixation filter converts raw gaze samples into the fixations and saccades that form the nodes and edges of the graph."},{"cited_title":"TSMixer: Lightweight MLP-mixer model for multivariate time se- ries forecasting,","cited_arxiv_id":null,"evidence_quote":"TSMixer serves as a state-of-the-art time-series baseline that the proposed model is compared against and outperforms."}],"review_version":1}