{"id":"9e43e7c8-5e0a-49f1-ac33-346c76eb65d4","arxiv_id":"2511.17581","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"EgoCogNav jointly predicts walking path, head motion, and perceived route uncertainty from egocentric sensors, and the CEN dataset makes such joint forecasting possible.","lead":"This paper presents a machine-learning model that uses first-person video, gaze, and motion to predict where a person will walk, where they will look, and how uncertain they feel about their route. It introduces a 6-hour egocentric navigation dataset in which people report their uncertainty in real time, aimed at assistive wayfinding and socially aware robots.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth uncertainty labels are confounded by instructions explicitly prompting scanning/confirmation behaviors at moments of self-reported uncertainty, so the behavioral-correlation claim may be protocol-induced.","rationale":"The reader's weakest assumption correctly identifies the continuous Xbox self-report as the empirical foundation for the cognitive claim. This stress-test refines that concern into a specific, load-bearing threat: the experimental protocol explicitly instructs participants to emit the very behaviors that are later used to validate the uncertainty-behavior coupling. This is not merely a statistical artifact; it is a demand characteristic that can manufacture the correlation the paper claims to discover. The central claim has two parts: joint forecasting (which is plausible) and behavioral correlation plus forecasting improvement (which depends on the validity of the uncertainty signal). If the self-report is unreliable or protocol-induced, the behavioral correlation results in Tables 2-3 and the qualitative 'alignment' claims lose their meaning. The paper currently provides no control for this, and the absence of error bars or per-participant analysis further weakens confidence. However, the dataset and task formalization are still valuable, and the motion forecasting improvements (Table 1) do not hinge on the cognitive ground truth as directly. A control study or a targeted re-analysis could settle the concern, so CONDITIONAL remains appropriate; the reader's verdict does not need to change, but the condition should be made explicit.","tokens_in":12369,"tokens_out":4834,"duration_ms":47617,"concrete_test":"Run a control condition with a separate group (or within-subject counterbalanced session) where participants navigate the same routes without the instruction to perform route-finding behaviors when uncertain—only self-report uncertainty. If the correlation between self-reported uncertainty and annotated scanning/confirmation/hesitation behaviors drops substantially relative to the original protocol, the behavioral-coupling claim is largely a task artifact. Alternatively, re-analyze the existing data by comparing the model's predicted-uncertainty elevation for non-instructed behaviors (e.g., WRONG, BACK) against instructed ones (SCAN, CONFIRM) and require the effect to persist for the non-instructed set.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 states participants were 'explicitly instructed to perform route-finding behaviors when uncertain such as scanning the surroundings for cues, confirming signage or landmarks.' This creates a direct demand characteristic: high uncertainty self-reports and the very behaviors used for validation (SCAN, CONFIRM, and arguably HES/LB) are co-produced by instruction, not independently coupled. Because the model is trained to regress these self-reports (Eq. 4) and then evaluated on behavioral correlation (Tables 2-3), the central claim that 'learned uncertainty strongly correlates with human-like behaviors' may be tautological. Additionally, the dual-task of holding and continuously adjusting a joystick while navigating can alter natural locomotion and decision-making, especially during hard wayfinding, biasing both the uncertainty labels and the behavior annotations. The paper provides no validation of the self-report (e.g., inter-rater reliability, comparison with retrospective ratings or physiological signals), and no control for the instruction-induced coupling. Without this, the empirical foundation for the cognition-aware aspect of the claim is insecure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes EgoCogNav, a multimodal egocentric navigation framework that jointly forecasts perceived path uncertainty, body-frame trajectory, and head motion from egocentric video, gaze, motion history, and navigation goals. The authors introduce the Cognition-aware Egocentric Navigation (CEN) dataset, containing 6 hours of real-world recordings from 17 participants across 42 indoor/outdoor sites, with continuous Xbox-controller self-reports of perceived uncertainty and annotations of behaviors and environment types. The architecture combines a frozen DINOv2 visual encoder, an action module for motion/gaze/goal encoding, a cognition head that regresses the self-reported uncertainty (Eq. 4), and adaptive goal conditioning (Eq. 1). Experiments compare against constant-velocity, linear-extrapolation, transformer, and EgoCast baselines, plus EMU/PATH-U uncertainty baselines, reporting improved trajectory/head forecasting and uncertainty–behavior correlations. Ablations and qualitative examples are also provided.","tokens_in":12625,"tokens_out":4190,"duration_ms":43286,"significance":"If the empirical claims hold, the CEN dataset would be a novel resource for cognition-aware egocentric navigation, and the task formulation—jointly predicting perceived uncertainty with motion—could enable assistive wayfinding and socially-aware navigation systems. The paper's strengths include the multimodal real-world data collection effort, the integration of cognitive state into a forecasting architecture, and the proposal of a concrete evaluation protocol with several metrics. However, the central claim that learned uncertainty 'strongly correlates with human-like behaviors' currently rests on a data-collection protocol that explicitly instructed participants to perform those very behaviors when uncertain, and on a supervised training signal derived from the same self-reports. The reported quantitative gains also lack error bars, seed counts, and split details, making it difficult to assess reliability. With additional validation and re-analysis, the work could become a solid contribution; in its present form the empirical foundation is insecure.","major_comments":[{"comment":"The data-collection protocol explicitly instructed participants to 'perform route-finding behaviors when uncertain such as scanning the surroundings for cues, confirming signage or landmarks.' This creates a demand characteristic that co-produces the self-reported uncertainty labels and the very behaviors used for validation (SCAN, CONFIRM, and arguably HES/LB). Because the uncertainty head is trained to regress those same self-reports (Eq. 4), the reported correlations in Table 3 (ΔU, effect sizes) may largely reflect protocol-induced co-occurrence rather than an emergent coupling between perceived uncertainty and behavior. The paper needs to address this directly: e.g., validate the self-report against independent measures (retrospective ratings, physiological signals), include a control condition without the instruction, report inter-rater reliability of behavior annotations, or at mi","section":"§4.1, Eq. (4), Table 3"},{"comment":"No error bars, number of seeds, or data-split details are reported. With 17 participants and 6 hours of data, differences such as ADE 0.14 vs 0.12 (Table 1) may be within noise. The 'High Uncertainty Scenarios' subset in Table 1 is described as 'top 20% highest uncertainty,' but the selection criterion is unspecified; if it is based on human labels, evaluation on this subset could introduce selection bias. The authors should specify the exact split (e.g., leave-participants-out vs. leave-environments-out), report results across multiple seeds with confidence intervals, and state how the high-uncertainty subset is defined and whether it uses ground-truth or predicted uncertainty.","section":"§5.1, Tables 1–4"},{"comment":"The claim that the cognitive signal 'improves trajectory and head-motion forecasting' is not supported by the ablations. Table 4 removes auxiliary losses and trajectory losses, but does not ablate the cognition head, the uncertainty loss (Eq. 4), or the adaptive goal conditioning (Eq. 1). Without an ablation that removes the uncertainty prediction/conditioning while keeping all other components, the contribution of the cognitive module to motion forecasting is untested. The paper should add such an ablation, or soften the claim accordingly.","section":"§3.2, §5.1.2, Table 4"},{"comment":"The Spearman ρ=0.64 reported for uncertainty prediction measures supervised fit accuracy, since the model is directly regressed to human self-reports via Eq. (4). The abstract's phrasing 'learns the perceived uncertainty that strongly correlates...' implies an emergent or discovered signal, but the evaluation is on the same dataset whose labels were used for training. This circularity should be acknowledged explicitly, and the authors should report performance under a stricter protocol, e.g., leave-one-participant-out, or compare against a model trained without uncertainty labels to demonstrate that the learned representation captures something beyond the supervision signal.","section":"§5.1.1, Table 2"}],"minor_comments":[{"comment":"Typo: 'past window T1 = 30 steps (3s) and T1 = 10 steps (1s)' — the second T1 should be T2.","section":"§3.1"},{"comment":"Minor language issues: 'We propose' should not be capitalized mid-sentence; 'dataset consisting 6 hours' is missing 'of'; 'videp' should be 'video'; 'Each design choice hels improve' should be 'helps improve.'","section":"§1, §4.1"},{"comment":"The text refers to 'the validation set' while Section 5 states all results are on a held-out test set. Please clarify the evaluation protocol and use consistent terminology.","section":"§5.1.2"},{"comment":"The EgoCast baseline is adapted 'with more layers' but no details are given. Specify the adaptation for reproducibility.","section":"§5.1.1"},{"comment":"The paper claims to introduce a dataset 'to facilitate research in the field,' but no dataset release URL or availability statement is provided. For a dataset contribution, this is essential.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and potentially useful problem, and the dataset is a valuable resource if released. However, the central behavioral-correlation claim is threatened by the instructed-behavior confound, and the quantitative evaluation lacks statistical rigor. I would require additional validation or reanalysis before considering acceptance. The lack of a dataset/code release, despite the dataset being a core contribution, is also a concern for verifiability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know up front. First, this is a real contribution: a new joint task (trajectory, head motion, and continuous perceived uncertainty from egocentric sensors) and a new 6-hour, 17-participant multimodal dataset that fills a genuine gap. Second, the paper's strongest claim—that learned uncertainty \"strongly correlates\" with scanning, hesitation, and backtracking—is much weaker than it looks, because the protocol told participants to perform exactly those behaviors whenever they felt uncertain. That is a textbook demand characteristic, and it means the behavioral-correlation numbers in Tables 2 and 3 may be partly an artifact of the instructions, not an emergent coupling.\n\nWhat is actually good: the task is new and well-motivated, and the CEN dataset—if released—would be a community asset. The architecture is a sensible composition of DINOv2 features, attention fusion, and parallel decoders; the ablations are internally consistent; the baselines are reasonable given that the closest prior works have no public code. The paper honestly reports failure cases and limitations, which I respect.\n\nThe soft spots, in rough order of seriousness.\n\nFirst and most damaging: the uncertainty self-reports are the ground truth for the uncertainty head (Eq. 4), and the behaviors used for validation (SCAN, CONFIRM, HES, LB) were explicitly instructed to be performed when uncertain. So the ρ=0.64 in Table 2 measures how well the model fits the labels, and the behavioral correlations in Table 3 are confounded by protocol. The paper offers no control for this, no inter-rater reliability on the self-report, and no comparison with a retrospective rating or physiological signal. This needs to be addressed before the cognitive claim can be taken at face value.\n\nSecond: no error bars, no number of seeds, no data-split details, and no code or dataset release. The \"High Uncertainty Scenarios\" subset in Table 1 is defined only as \"top 20% highest uncertainty\" without specifying whether that is by human labels or model predictions. Fixable in revision, but exactly the details a referee needs.\n\nThird: framing. Calling the uncertainty head a \"latent state\" overstates what is a supervised regression on self-reports. That doesn't kill the contribution, but it should be described accurately.\n\nWho is this for: people working on egocentric navigation, human trajectory forecasting, and assistive wayfinding. The dataset and task definition alone justify a serious look. My recommendation: send it to peer review, but the referee should require a substantially revised version that addresses the instruction-effect confound, reports statistics properly, and releases the data or at least a thorough protocol analysis. If those fixes come through, this could be a solid addition to the literature.","headline":"A useful dataset-and-task paper whose cognitive-behavioral claim is entangled with its own data-collection instructions; worth a careful revise-and-resubmit, not a desk reject.","tokens_in":13119,"tokens_out":2941,"would_cite":true,"duration_ms":25762,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A person's moment-to-moment feeling of being lost can be predicted from egocentric video, gaze, and motion, and the predicted uncertainty improves trajectory and head-motion forecasts.","keywords":["egocentric navigation","perceived path uncertainty","trajectory forecasting","head-motion forecasting","multimodal learning","wayfinding","cognitive modeling","egocentric dataset"],"falsifier":"Run the same collection protocol with uncertainty measured passively (pupil dilation, gaze entropy, or heart-rate variability) instead of by joystick, and retrain or retest EgoCogNav; if its learned uncertainty no longer tracks behavioral events with a Spearman ρ near 0.64 and instead falls to the 0.13–0.17 range of the theory baselines, the model was reconstructing an instruction artifact rather than a cognitive state. A simpler check: hold out joystick labels for entire participants and see whether predicted uncertainty still predicts their hesitation and backtracking.","tokens_in":12271,"feed_emoji":"🧭","tokens_out":11070,"duration_ms":93907,"temperature":0.7,"pith_summary":"This paper tries to establish that the feeling of not knowing which way to go—perceived path uncertainty—can be learned as a moment-to-moment state from an egocentric stream of video, gaze, head motion, and recent walking, and that forecasting this state alongside trajectories and head pose yields better short-horizon predictions than ignoring cognition. The authors built EgoCogNav, which fuses these signals into a shared representation, predicts a single uncertainty value in [0,1], and uses that value to modulate how much the decoder trusts goal information versus raw sensory evidence. They also collected the CEN dataset: six hours of real indoor and outdoor navigation from 17 participants who continuously reported their uncertainty through a handheld controller. On held-out environments, the model's uncertainty estimates correlate with human reports much more strongly than entropy-theory or rule-based baselines, and predicted uncertainty rises at moments of scanning, hesitation, backtracking, and wrong turns. A sympathetic reader would care because current trajectory predictors treat navigation as homogeneous motion history; if this approach holds, assistive wayfinding and social robots could respond to confusion before it becomes visible in gait.","feed_headline":"A model predicts when walkers feel lost from video, gaze, and gait","feed_subtitle":"Uncertainty self-reports train a model to spot hesitation, backtracking, and wrong turns before they happen.","key_machinery":"The central mechanism is the uncertainty-gated fusion loop: a cognition head maps a time-pooled multimodal representation into a scalar belief U_t ∈ [0,1], and this belief modulates the feature mix through adaptive goal conditioning—h_tilde = (1 − U_t)·h_fuse + U_t·h_goal—so that when uncertainty is high, the decoder weights goal-consistent evidence more heavily. Parallel temporal decoders then forecast the body-frame trajectory (3-DOF: displacement and heading change) and head rotations (6-DOF), while auxiliary classifiers for environment type and behavioral events regularize the shared backbone during training. What makes the cognitive state learnable is the dataset's ground-truth anchor:","core_discovery":"EgoCogNav jointly forecasts a body-frame trajectory, head rotations, and a scalar perceived uncertainty from 3 seconds of egocentric video, body motion, head pose, gaze, and a navigation goal. A cognition head pools these features into U_t ∈ [0,1], trained to match human self-reports, and this uncertainty gates the fused features so high uncertainty relies more on goal-consistent evidence. On held-out environments it outperforms the strongest baseline on trajectory and head-rotation error. Its uncertainty predictions reach Spearman ρ = 0.64, far above two non-learning baselines, and peak for wrong turns and backtracks while rising for hesitation and scanning. The paper reads this as evidence","pith_inferences":["Editorial inference: if the uncertainty signal transfers across people and places, it can be inverted into an environmental-difficulty heatmap: a few walk-throughs could reveal which junctions, occlusions, or sign placements reliably induce confusion, giving designers a continuous, cheap alternative to surveys.","Editorial inference: the goal-conditioning rule is a general recipe for embodied agents—use an internal confidence estimate to decide when to trust a goal prior rather than raw sensory evidence—and could be tested independently in simulated wayfinding agents with synthetic uncertainty labels.","Editorial inference: a sharper test of the cognitive claim would hold out joystick reports for entire participants and ask whether the model still predicts their behavior; if it does, the uncertainty head is recovering a real behavioral signal; if not, it is partly memorizing report styles.","Editorial inference: because the model outputs a single best path, its correlation with backtracking could reflect the model recognizing the same ambiguous scene rather than truly representing a decision conflict; a multi-hypothesis variant would separate these explanations."],"forward_implications":["A deployed system could flag high-predicted-uncertainty moments as prompts for assistance; the model's top-20% uncertainty moments match labeled difficulty events 45% of the time, versus 18–20% for the baselines.","Adding the cognitive signal changes motion forecasts, not just labels: on the hardest 20% of moments, EgoCogNav cuts ADE from 0.15 to 0.13 and head-rotation error from 0.088 to 0.083 against the strongest baseline.","Predicted uncertainty rises before route-finding behaviors—hesitation, scanning, wrong turns, backtracking—so the same model can be used to label environments by how confusing they are to navigate, not just by their geometry.","The CEN dataset gives the field a public, multimodal benchmark with synchronized video, gaze, head pose, trajectory, and moment-to-moment uncertainty reports, which the paper argues the closest prior egocentric navigation datasets do not provide.","The ablations show that both the horizon-weighted trajectory loss and the auxiliary environment/behavior classifiers contribute, meaning the improvements come from the full perception–cognition loop rather than any single input modality."],"fun_headline_variants":["From egocentric video to feeling lost: a model predicts walker uncertainty","A model reads doubt in your walk: predicts wrong turns from video and gaze","Predicting perceived uncertainty: when walkers hesitate, scan, backtrack","EgoCogNav: forecasting when walkers doubt their path"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claim depends on the 17 participants' continuous handheld-controller self-reports of perceived uncertainty being a faithful, non-reactive measure of a real internal state—and on the instructed route-finding behaviors (scan, confirm) not manufacturing the very uncertainty–behavior correlation the paper reports.","fun_headline_variants_meta":{"raw":{"variants":["From egocentric video to feeling lost: a model predicts walker uncertainty","A model reads doubt in your walk: predicts wrong turns from video and gaze","Predicting perceived uncertainty: when walkers hesitate, scan, backtrack","EgoCogNav: forecasting when walkers doubt their path"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000646,"raw_usage":{"total_tokens":2774,"prompt_tokens":686,"completion_tokens":2088,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":430,"completion_tokens_details":{"reasoning_tokens":2010}},"tokens_in":430,"tokens_out":2088,"duration_ms":14632,"temperature":1.0,"reasoning_tokens":2010,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T22:02:54.135871+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same collection protocol with uncertainty measured passively (pupil dilation, gaze entropy, or heart-rate variability) instead of by joystick, and retrain or retest EgoCogNav; if its learned uncertainty no longer tracks behavioral events with a Spearman ρ near 0.64 and instead falls to the 0.13–0.17 range of the theory baselines, the model was reconstructing an instruction artifact rather than a cognitive state. A simpler check: hold out joystick labels for entire participants and see whether predicted uncertainty still predicts their hesitation and backtracking.","supporting_citations":[],"review_version":1}