{"id":"9305aa1f-18a0-4f94-b227-3196c158d78f","arxiv_id":"2607.08083","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Under high perceptual load, routing brief probes to the channel with higher HeadRoom-estimated availability cuts response time versus the less available channel.","lead":"HeadRoom is a tiny on-device model that watches what you see and hear through smart glasses and guesses which sense has spare capacity right now. Under heavy load, sending a notification to the freer channel cut reaction time versus the busier one.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No stronger load-bearing flaw than the reader's proxy concern; Video-3 Model-vs-Inverse effect and trial-level correlations still stand as reported.","rationale":"The reader's weakest-assumption diagnosis is exactly the softest link in the argument chain: everything downstream (routing rule, Video-3 effect, abstract claim) inherits whatever validity the prediction-error proxy possesses. The paper already reports the key positive evidence (condition effect + continuous correlations) and an honest limitations section that flags the passive-probe task, the lack of Model>Random advantage, and the single high-load video. No internal inconsistency, statistical error, or hidden circularity appears that would push the verdict below CONDITIONAL; the edge-footprint numbers and open models further support the systems contribution. Hence the existing CONDITIONAL / HIGH-confidence judgment is left unchanged. The concrete test above isolates whether the proxy's behavioral relevance survives removal of the selection filter and low-level confounds—the single check that would most cleanly confirm or refute the load-bearing assumption.","tokens_in":12873,"tokens_out":645,"duration_ms":37775,"concrete_test":"Recompute the Video-3 trial-level Pearson/Spearman correlations and the channel-adjusted LME after (a) removing the model-based probe-time filter (use all 40 uniform bins) and (b) residualizing availability against the hand-crafted features in Table 7 (motion energy, spectral flux, etc.). If either the Model–Inverse β or the Δavail–RT correlation loses significance (p>0.05) or halves in magnitude, the proxy-validation claim weakens materially.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (high-load routing to the model's higher-availability channel yields lower RT than the lower-availability channel) rests on next-frame/window prediction error (frozen MobileNetV3-Small 576-d MLP for vision; 31-d MFCC+flux MLP for audio), after per-sequence EMA z-score normalization and inversion (§3.1–3.2, Alg. 1), being a faithful real-time proxy for residual channel capacity. The paper supplies supporting evidence—Video 3 Model–Inverse ΔRT ≈ −114 ms (p=.021), LME β≈−0.13–0.18 on log-RT, and trial-level r(avail, RT)≈−0.11 / r(Δavail, RT)≈−0.20—but that evidence is confined to one held-out high-demand clip, between-subjects cells of n=7–8, probe times pre-filtered by the same model (max α≥0.3), and a pure detection task with fixed-location visual probes. If the proxy mainly tracks low-level motion/energy rather than true residual capacity, or if the filter + small-n design inflates the ranking effect, both the RT difference and the correlations lose their interpretation as validation of channel-aware routing.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"HeadRoom estimates moment-to-moment visual and auditory channel availability from egocentric streams by treating next-step prediction error (frozen MobileNetV3-Small 576-d embeddings + MLP for vision; 31-d MFCC/RMS/flux/ZCR + MLP for audio) as a proxy for residual perceptual capacity, after online EMA z-score normalization and inversion (§3.1–3.2, Alg. 1). A routing module then selects the higher-availability channel (with occupancy threshold τ=0.3 and tie δ=0.05). A controlled probe-detection study (N=25, analyzed N=22) on three held-out Aria scenarios under Model / Inverse / Random routing finds no reliable pooled effect, but under the highest-demand Video 3, Model yields faster RT than Inverse (≈−114 ms, p=.021; LME β≈−0.13 to −0.18 on log-RT), with trial-level correlations between availability (and availability delta) and RT. Edge feasibility is shown on Meta Quest 3S (mean ~11 ms/step, 0.625 MB).","tokens_in":13258,"tokens_out":1418,"duration_ms":19647,"significance":"If the high-load result generalizes, the work supplies a practical, open-source, edge-deployable primitive for modality-aware notification routing on wearables and XR devices—addressing a real gap between Multiple Resource Theory and deployable systems. Strengths include the lightweight self-supervised design, explicit runtime/memory numbers on commodity XR hardware, open models/ONNX/Unity prototype, and a psychophysical evaluation with mixed-effects models and availability–RT correlations rather than only subjective load. The contribution is scoped carefully to detection under high demand and does not overclaim ecological superiority over random routing.","major_comments":[{"comment":"§5.2.1–5.2.2 and Tables 3–5: The central claim is supported only in Video 3 (between-subjects cells n=7–8). Pooled Model vs Inverse is non-significant (Wilcoxon p=.166; paired t p=.435), and Model is not significantly faster than Random in Video 3 (p=.209; LME still favors both over Inverse). With counterbalancing producing small cells and only one high-demand clip showing the effect, the evidence for adaptive routing as a general design primitive is thin. Either power the high-demand contrast adequately (within-subjects or larger N), pre-register the Video-3 focus, or substantially qualify the abstract/conclusion claim that currently reads more broadly than the data.","section":null},{"comment":"§5.1 / Appendix A.2 (probe selection): Probe onsets are chosen only at moments where the model’s higher availability ≥0.3. This couples the experimental stimulus set to the same signal under test and can inflate Model–Inverse ranking differences by excluding low-confidence or near-tie moments. Report sensitivity analyses without the filter (or with τ varied), and clarify how many candidate bins were discarded; otherwise the RT difference and availability–RT correlations partly reflect selection rather than pure routing validity.","section":null},{"comment":"§3.1–3.2 and §5.2.3: The load-bearing assumption is that next-frame/window prediction error (after normalization) indexes residual channel capacity. Supporting correlations with low-level features are weak (Appendix Table 7; strongest r=−0.309 motion vs visual availability), and trial-level r(avail, RT)≈−0.11 / r(Δavail, RT)≈−0.20 are modest. The paper needs either (a) a stronger external validation of the proxy (e.g., against dual-task cost or known load manipulations independent of the routing labels) or (b) explicit framing that the result validates relative ranking under this operationalization, not that prediction error equals true spare capacity. Without that, the interpretation of Model vs Inverse as ‘channel-aware routing’ remains under-constrained.","section":null},{"comment":"§5.2 and Limitations: Visual probes are systematically faster than auditory ones (574 vs 636 ms), which the authors attribute to visual priming from continuous screen viewing. Combined with fixed upper-right probe location and pure detection (not interpretation/action), this limits claims about real notification routing. The manuscript already notes ecological limits; the major issue is that the abstract and contribution statements still present the result as evidence for adaptive notification routing in wearables. Tighten those statements to match the detection-task, high-load, Model-vs-Inverse scope, or add a richer secondary measure.","section":null}],"minor_comments":[{"comment":"Abstract and §1: N=25 is stated; analyzed N=22 after exclusion for the 80% training threshold. Report analyzed N consistently in the abstract.","section":null},{"comment":"Eq. (1) and Appendix A.1: τ=0.3 and δ=0.05 are called ‘practical operating values’; post-hoc note that benefits weaken at τ≤0.2 is useful—move a brief sensitivity statement into the main text so readers see parameter dependence without only the appendix.","section":null},{"comment":"Figure 3 / Tables 3–5: Provide error bars or CIs on the per-condition means and state whether means are participant-level or trial-level aggregates.","section":null},{"comment":"§4.1: Live Aria streaming observations are qualitative and pointed to the website; a short quantitative summary (e.g., direction of availability shifts under abrupt sound vs high motion) in the main text would strengthen the feasibility narrative.","section":null},{"comment":"Typos/clarity: ‘wihtout’ (§2), ‘Promemassist’ vs ‘Promemas-sist’ inconsistency with citation [22], and ‘either when the difference is small’ in the routing description could be tightened.","section":null},{"comment":"NASA-TLX null result is informative; report effect sizes or Bayes factors if space allows so readers can judge evidence of absence vs underpowering.","section":null}],"recommendation":"major_revision","confidential_remarks":"Fit for a solid HCI/systems venue is reasonable if the authors narrow claims and address the probe-filter and small-cell issues. The open-source edge prototype is a genuine plus. I would not reject on novelty grounds; the predictive-coding operationalization is a legitimate design choice even if the proxy remains imperfect. Main risk is over-generalization from one high-load clip and a detection task."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean HCI systems paper that actually ships something usable. What is new is the concrete dual-predictor pipeline: frozen MobileNetV3 embeddings + tiny MLP for vision, 31-d MFCC/flux MLP for audio, online EMA z-score normalization, inverted to availability, then a simple routing rule. They get it under 15 ms and <1 MB on Quest 3S, open-source the ONNX bits, and show a live Aria stream. That is real engineering value for wearable/XR people who need modality choice without a cloud round-trip.\n\nThe behavioral result is narrower but honest. Pooled Model vs Inverse is null; Model is not reliably better than Random. Only Video 3 (high demand, comprehension at 50%) shows the effect: Model ~114 ms faster than Inverse, large effect, mixed models on log RT hold after channel adjustment, and trial-level availability/delta-availability correlate with RT in the expected direction. Detection rates stay high; NASA-TLX does not move. The limitations section already flags the passive probe task, fixed visual location, and the fact that this is detection not deeper processing. So the abstract’s “under high perceptual load” claim matches the data; the broader framing does not.\n\nSoft spots in proportion: the predictive-coding proxy is the load-bearing assumption and is only weakly checked against low-level features (strongest r = –0.3 with motion). Probe times are pre-filtered by the model itself (α ≥ 0.3), cells are n≈7–8 between subjects, and three people were dropped for the 80% threshold. Free parameters (τ, δ, warm-up, log compression) are set a priori and not stress-tested much. None of that invents the Video-3 contrast or the correlations, but it does keep the result conditional. Circularity is low: prediction error is independent of the RT measure.\n\nMath and stats are standard and readable; citations sit properly on MRT, interruption timing, and predictive coding without overclaiming. Who it is for: anyone building adaptive notifications on glasses or XR who wants a runnable baseline and a carefully scoped existence proof. I would bring it to reading group for the systems piece and the honest boundary conditions. It deserves a serious referee rather than a desk reject; expect revision pressure on ecological validity and proxy validation, not rejection of the core result.","headline":"Solid lightweight systems paper with a real high-load RT effect and open edge code; the proxy and scope are the soft spots, not a collapse of the claim.","tokens_in":13805,"tokens_out":573,"would_cite":true,"duration_ms":6226,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Under high perceptual load, routing wearable notifications to the sensory channel with more residual capacity reduces response time.","keywords":["multimodal interaction","wearables","disruption","channel availability","notification routing","egocentric sensing","edge deployment","predictive coding"],"falsifier":"In a high-demand egocentric scene, if participants respond equally fast or faster when probes are deliberately sent to the higher-prediction-error channel than to the lower-prediction-error channel, the central routing claim is false.","tokens_in":13769,"feed_emoji":"👓","tokens_out":709,"duration_ms":19221,"temperature":0.7,"pith_summary":"Wearables such as smart glasses can deliver notifications by sight or by sound, yet designers still lack a practical way to pick the right channel at the right moment. HeadRoom is a lightweight, edge-deployable pipeline that watches egocentric video and audio and estimates residual capacity in each sensory channel by treating next-step prediction error as a real-time proxy for load. In a controlled study, when overall perceptual demand was high, sending brief probes to the channel the model judged more available produced faster responses than sending them to the less available channel; under lower demand the difference largely disappeared. The full pipeline runs on ordinary wearable hardware in about 11 ms per step with a sub-megabyte model footprint. If the approach generalizes, future wearable and immersive systems could interrupt people less disruptively by matching output modality to moment-to-moment channel headroom.","feed_headline":"Route alerts to the freer channel; RT drops under load","feed_subtitle":"Tiny on-device predictors estimate visual and auditory headroom from egocentric video and audio.","key_machinery":"Channel availability, defined as one minus the normalized next-step prediction error of a frozen-MobileNet embedding MLP (vision) and a 31-dimensional MFCC/flux MLP (audio). Higher prediction error is treated as higher occupancy; a simple comparison of the two availability scores decides the routing target.","core_discovery":"HeadRoom demonstrates that visual and auditory channel availability—estimated continuously as inverted, online-normalized prediction error of separate lightweight next-step predictors on egocentric streams—can be used for adaptive notification routing. Under high perceptual load, routing probes to the more available channel measurably reduces response time relative to routing them to the less available channel.","pith_inferences":["The same prediction-error idea could be extended to haptics, giving a third routing option when both vision and audition are occupied.","Availability traces might serve as a continuous, sensor-free secondary measure of channel load in dual-task psychology experiments.","Whether the millisecond-scale detection gains survive for richer notifications that require interpretation or action remains an open, testable question."],"forward_implications":["When perceptual demand is high, systems can reduce response cost by preferring the currently freer sensory channel.","The same continuous availability signal can also help decide when to interrupt, not only which modality to use.","Sub-megabyte models and ~11 ms on-device latency make real-time channel-aware routing practical on contemporary XR headsets.","Routing benefits are largest under elevated demand; under low demand the advantage shrinks toward chance."],"fun_headline_variants":["Freer-channel routing cuts RT under high perceptual load","Edge headroom predictors route alerts to freer sensory channel","Route probes to more available channel; RT drops under load","Online visual/auditory headroom estimates enable adaptive routing","HeadRoom freer-channel alerts reduce response time when load is high"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That how poorly a simple next-moment predictor anticipates the next visual embedding or audio feature vector is a faithful real-time stand-in for residual capacity in that sensory channel.","fun_headline_variants_meta":{"raw":{"variants":["Freer-channel routing cuts RT under high perceptual load","Edge headroom predictors route alerts to freer sensory channel","Route probes to more available channel; RT drops under load","Online visual/auditory headroom estimates enable adaptive routing","HeadRoom freer-channel alerts reduce response time when load is high"]},"model":"grok-4.5","effort":"low","cost_usd":0.007982,"raw_usage":{"total_tokens":1819,"prompt_tokens":637,"num_sources_used":0,"completion_tokens":86,"cost_in_usd_ticks":79820000,"prompt_tokens_details":{"text_tokens":637,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1096,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":637,"tokens_out":86,"duration_ms":10303,"temperature":1.0,"reasoning_tokens":1096,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T13:23:06.790523+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"In a high-demand egocentric scene, if participants respond equally fast or faster when probes are deliberately sent to the higher-prediction-error channel than to the lower-prediction-error channel, the central routing claim is false.","supporting_citations":[],"review_version":1}